DocsQuick StartAI News
AI NewsAI Grading Homework: Why Are the Scores Inconsistent?
Industry News

AI Grading Homework: Why Are the Scores Inconsistent?

2026-08-24T16:03:49.107Z
AI Grading Homework: Why Are the Scores Inconsistent?

The latest research from Cardiff University and the University of Melbourne shows that when ChatGPT grades undergraduate biology essays, its score for a single paper can differ from a human grader’s by as much as 40 points, with a clear tendency to mark down high-scoring papers and mark up low-scoring ones. Large language models are well suited to serving as feedback assistants, but they remain far from being able to reliably take responsibility for grading.

Why Are AI-Graded Assignment Scores Unstable?

On August 24, 2026, a new study by Cardiff University and the University of Melbourne poured cold water on educational AI: generative AI can read a batch of essays quickly, but it still cannot assign grades as consistently, reproducibly, and explainably as a teacher.

According to the research findings disclosed by Cardiff University, the researchers asked two versions of ChatGPT to evaluate 50 undergraduate biology essays, scoring them according to seven criteria and using four different prompting methods. The research team then compared the model-generated scores with human scores, examining not only the average total scores but also fluctuations across different criteria and essays.

The conclusion was not encouraging: although the overall average score given by AI was sometimes close to the human-assigned average, the discrepancy quickly widened when it came to individual essays and specific scoring dimensions. Except in one case, the model’s average scores were generally higher than the teachers’ scores. At the level of average scores, the largest gap between AI and human grading reached 16.1 points; for individual essays, the maximum discrepancy reached 40 points.

More notably, the model displayed a typical tendency to “regress toward the mean”: high-scoring essays tended to be marked down, while low-scoring essays tended to be marked up, causing a large number of scores to cluster in the middle range.

This is not simply a matter of “AI occasionally assigning the wrong score.” It is a more difficult problem: when grading tasks are subjective, context-dependent, and open-ended, large language models may be able to generate feedback that appears reasonable, yet still be unable to apply the same scoring scale consistently.

A teacher grading student essays at a computer, with a comparison of AI and human scores displayed on one side of the screen

A Similar Overall Score Does Not Mean Reliable Grading

When educational institutions introduce automated grading systems, they are most easily persuaded by one metric: the model’s average score for a class is close to the teacher’s average.

But this study reminds us that averages often conceal the real risks.

Suppose a teacher gives a group of essays an average score of 72, while AI gives them an average of 74. Judging only from class-level statistics, it might seem reasonable to conclude that the model is “basically reliable.” But if one essay receives a 90 from the teacher and only a 70 from the model, while another receives a 50 from the teacher and a 70 from the model, the two averages may still be close—even though the outcomes for the individual students are entirely different.

The value of grading is not merely to produce roughly similar scores for the class as a whole. It must answer three more specific questions:

  • What level has this assignment actually reached?
  • On which scoring dimension did it lose points?
  • If the student requests a review, can the teacher explain why this score was given?

In open-ended assignments such as essays, compositions, design reports, and experimental analyses, grades are essentially a comprehensive assessment of the quality of argumentation, conceptual understanding, use of evidence, structural organization, and original expression—not a simple count of words, grammatical errors, or keyword occurrences. Differences may naturally exist among teachers, but teachers are generally able to calibrate their grading in light of course objectives, classroom progress, and students’ backgrounds.

The problem with large language models is that they can imitate the linguistic form of such judgments without necessarily possessing a stable internal scale. In other words, they are very good at explaining “why this essay is good,” but may not be able to convert “good” into the same score every time.

Change the Prompt, and the Results May Change Too

The researchers used four different prompting methods, and this aspect of the design is crucial. It tested not whether the model could “grade essays,” but whether it could maintain consistent grading standards when the task description changed.

In actual use, a teacher might write prompts such as:

  • Please assign a total score and provide feedback according to the grading criteria;
  • Analyze each criterion first, then calculate the total score;
  • Grade according to strict university-level teaching standards;
  • Focus on argumentation, evidence, and originality rather than language alone.

To humans, these instructions may seem broadly equivalent. To a large language model, however, they may imply different allocations of attention. If one prompt emphasizes linguistic accuracy, the model may place greater weight on grammar and expression; if another emphasizes argumentative structure, the model may assign greater importance to the completeness of the content.

Model outputs may also be affected by the order of the context, sample answers, system instructions, temperature settings, and descriptions of a particular essay earlier in the conversation. Even when the same model and the same grading criteria are used, changing the input format may still produce different results.

This is also what distinguishes large language model grading from traditional rule-based automated grading. Rule-based systems are rigid, but their behavioral boundaries are relatively clear. Large language models have stronger language-understanding and feedback capabilities, yet may display opaque flexibility in borderline cases.

In educational settings, such flexibility is not necessarily an advantage. Students need evaluation rules that are relatively stable, not polished but inconsistent feedback. If changing the prompt can add 10 points to the same essay, the model will struggle to serve as the final arbiter in high-stakes educational decisions.

AI May Be Better at “Identifying Problems” Than at “Determining Scores”

This does not mean that AI-assisted grading has no value. On the contrary, the most valuable applications of large models in education may not involve directly replacing the teacher’s final judgment, but rather handling the repetitive, low-risk tasks in the grading process first.

For example, AI can help teachers:

  1. Quickly flag obvious problems with grammar, formatting, and citations;
  2. Organize an essay’s strengths and weaknesses according to the grading criteria;
  3. Identify places where evidence is missing, concepts jump abruptly, or statements contradict one another;
  4. Generate revision suggestions for students at different levels;
  5. Summarize common problems across a batch of assignments to help teachers design the next lesson;
  6. Provide multiple revision directions for the same piece of writing rather than outputting a single score.

These tasks are closer to “retrieving, summarizing, and generating feedback,” and therefore allow greater tolerance for error. Even if AI overlooks a problem, the teacher can still make the final judgment.

However, the situation is different when scores affect scholarships, course completion, admission to further education, graduation, or academic evaluation. In such cases, the grade is not merely a teaching prompt but a formal decision. Model bias, instability, and insufficient explainability can all translate directly into tangible losses for students.

Long-form writing such as essays may be particularly susceptible to the influence of writing style, linguistic background, and writing habits. Fluent prose and dense terminology do not necessarily indicate correct ideas; plain expression does not mean weak argumentation. If a model more readily rewards text that “looks like a standard answer,” it may steer students toward formulaic writing and, in turn, suppress genuinely individual expression.

“Clustering Toward the Middle” Is Particularly Dangerous for Educational Assessment

The “pulling down high scores and raising low scores” observed in the study has another easily overlooked consequence: it weakens the differentiation between students’ performance levels.

To a teacher, an essay that is close to full marks and one that barely passes often represent entirely different learning outcomes. The former may demonstrate solid mastery of knowledge and independent analytical ability, while the latter may merely fulfill the basic requirements. If AI tends to assign both essays scores in the middle range, it may appear to reduce extreme errors, but in reality it erases the differences among students that are most educationally meaningful.

This phenomenon may also affect teachers’ understanding of students’ learning. The real problem in a class may be that “a small number of students have mastered the material, while others have not understood it at all,” but the score distribution generated by AI may suggest that “everyone is more or less at the same level.” If teachers adjust their instruction based on that result, they may overlook the students who genuinely need focused support.

From the model’s perspective, this tendency is not difficult to understand. During training, large language models encounter vast quantities of linguistic data and are generally better at generating answers that conform to common patterns, social expectations, and middle-range standards. When faced with an exceptionally strong or weak open-ended text, the model may lack a sufficiently stable frame of reference and therefore tend to produce a middle-ground judgment that “sounds reasonable.”

The problem is that educational assessment often needs to identify precisely such extremes: genuine breakthroughs, serious misunderstandings, unique but defensible perspectives, and creativity that is not yet fully formed but deserves to be preserved.

This Study Does Not Prove That “AI Cannot Grade at All”

It is important to note that the study involved 50 undergraduate biology essays, tested two versions of ChatGPT, and used specific grading criteria and prompting methods. The findings therefore cannot be simply generalized to every discipline, model, or educational system.

Tasks with a higher degree of standardization—such as multiple-choice questions, calculation problems, vocabulary dictation, format checking, and some factual questions—are generally more suitable for automated processing. For compositions and essays, model performance may also improve significantly if the grading criteria are sufficiently clear, the training data are ample, and multiple rounds of human calibration are conducted.

But the value of the study lies precisely in moving the discussion from “Can a model generate grading feedback?” to “Can a model reliably assume responsibility for assigning scores?” These are not the same thing.

In the past, many product demonstrations showed only a single successful case: AI read an essay, generated detailed feedback, and then gave a seemingly reasonable score. Once deployed in a school, however, the system must handle thousands of assignments with different levels, styles, and backgrounds, while also standing up to reviews, appeals, and comparisons across teachers. A single sufficiently good output does not mean that the system will remain stable over long-term operation.

A More Realistic Deployment Model: Let AI Serve as a Copilot

If schools or teachers still wish to use large models to grade assignments, the more prudent approach is not to remove teachers from the process, but to establish a tiered human–machine collaboration mechanism.

The following practices may be considered:

  • Let AI conduct preliminary grading, while teachers assign the final scores. The model can output evidence for each criterion, points of concern, and a suggested score, but should not write directly to the official gradebook.
  • Break the total score down into auditable dimensions. Each subscore should require the model to cite specific passages from the essay rather than merely provide a conclusion.
  • Set thresholds for human review. Assignments with large discrepancies from historical teacher scores, low confidence, or potentially original viewpoints should automatically be referred to a human.
  • Use multiple runs and cross-checks among different models. If results vary significantly across prompts or models, they should not be adopted directly.
  • Conduct regular calibration using institutional data. It is not enough to demonstrate reliability using a vendor’s general evaluation results. Consistency should be tested with anonymized assignments from the specific course, discipline, and semester.
  • Preserve students’ rights to appeal and request review. AI outputs should be traceable recommendations, not black-box conclusions that students have no way to challenge.
  • Define clear boundaries for data and privacy. Student assignments constitute educational data. Uploading them to external AI services without clear notice and consent may raise issues involving privacy, intellectual property, and data compliance.

A truly mature system should allow teachers to inspect the model’s basis for judgment, modify scores, record the reasons for review, and feed human corrections back into course-level quality monitoring. AI’s role should be more like that of a tireless teaching assistant—not the chief examiner holding the final grade sheet.

The Key Metrics for Educational AI Should Not Be Limited to Efficiency

The appeal of automated grading is obvious: it can process in minutes the preliminary screening and feedback that previously required teachers to spend hours completing. But efficiency is only the first layer of metrics for educational technology.

More important questions include:

  • When the same assignment is graded repeatedly, are the results stable?
  • Are students from different linguistic backgrounds, with different writing styles and from different fields, affected differently?
  • Can the model distinguish factual errors, immature expression, and deliberate innovation?
  • Can teachers and students understand the grading criteria?
  • Do students know how their assignments are processed and how long their data will be retained?
  • When a dispute occurs, can the school reconstruct the basis on which the model made its judgment?

If these questions have no answers, then “a tenfold increase in grading speed” may not represent technological progress. It may simply replace the previously visible biases of human processes with automated biases that are more difficult to hold accountable.

As of today, the signal from this study is clear: large language models are already capable enough to participate in the grading process, but they are not yet stable enough to independently determine grades for complex written assignments. What educational institutions should truly pursue is not having AI grade every assignment in place of teachers, but freeing teachers from mechanical labor so that they can devote their time to what models are hardest to replace—understanding students, evaluating complex ideas, and explaining why an answer merits a particular assessment.

In this sense, the endpoint of AI-assisted grading may not be “unmanned assessment,” but rather “machines process the evidence, while teachers remain responsible for the judgment.”

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: