AI & Society

Even Teachers Deferred to AI When the Grade Was Too Harsh

A teacher holding a marked 50/50 exam paper considers an AI assistant pointing to a student’s answer.
An AI assistant points to an exam response as a teacher reconsiders a tough marking decision.

AI grading made teachers more willing to accept an unfairly low mark. In a preregistered experiment involving more than 1,300 Greek teachers, participants saw identical student work and the same recommended score of five out of ten. When the harsh recommendation was labelled as an algorithm’s, teachers judged it 22 per cent fairer than when it was attributed to a human colleague.

The system did not analyse the work or produce a genuine score. Researchers varied only the label attached to the recommendation. That makes the result especially revealing: the belief that advice came from an algorithm changed professional judgement even though the evidence, grade and checklist remained identical.

The AI grading experiment used the same work twice

The study recruited Greek primary and secondary school teachers for an online experiment. Participants reviewed a short student response using an explicit grading checklist. They were randomly shown work designed to deserve either eight out of ten or two out of ten, while the recommendation displayed beside it was always five.

For the strong answer, five was unfairly harsh. For the weak answer, five was unfairly generous. Half the teachers were told the recommendation came from a human colleague and half were told it came from an algorithm. Because everything else was held constant, differences in perceived fairness could be attributed to the source label rather than the student’s performance.

The reported sample contained 677 teachers in the harsh condition and 662 in the lenient condition, for a total of 1,339. The study was randomised and preregistered, and the research materials and data record provide a route for other researchers to inspect or reproduce the analysis.

Teachers gave harsh algorithmic advice more credit

In the harsh condition, teachers rated the discrepancy as more acceptable when the recommendation carried the algorithm label. The average fairness-gap measure was 1.584 for the algorithm and 1.384 for the human colleague. With controls, the estimated difference was 0.300 and statistically significant, which the authors describe as a 22 per cent relative increase.

The result did not repeat symmetrically when the suggested grade was too generous. Differences between the algorithm and human-labelled recommendations in the lenient condition were not statistically significant. Teachers were not automatically persuaded by every machine score. Their deference appeared when the system imposed a harsher outcome.

That asymmetry matters. A general preference for algorithms might produce similar effects in both directions. Instead, the experiment suggests that perceived objectivity interacts with a punitive recommendation. A low machine score may feel like a neutral measurement, while an equally low human score is more readily seen as judgement that can be questioned.

What AI grading means for students and teachers

A student should never lose marks solely because a teacher assumes an algorithm is more objective than a colleague. Schools using automated recommendations need a documented human review that starts from the student’s work and rubric, not from an attempt to justify the number already on screen. Reviewers should record why they agreed or disagreed before the final grade is issued.

Teachers need access to the model’s purpose, evidence and known error patterns. A score without an explanation can create authority without understanding. If an automated system was trained on past marking, educators should know which subjects, age groups and languages were represented, and whether the tool was validated for the exact assessment now being graded.

Students and parents also need a meaningful appeal route. Human review should involve someone able to change the outcome, not merely explain the algorithm’s recommendation. The institution should retain the original work, rubric, model output and final reasoning so that an unfair mark can be reconstructed rather than treated as an untraceable technical event.

Why a machine can look stricter and fairer

The researchers tested whether perceived ability and responsibility helped explain the effect. Their mediation analysis suggested that teachers saw the algorithm as more capable and less personally responsible for the harsh outcome, and those perceptions accounted for more than half of the difference. The authors appropriately treat that mechanism as suggestive rather than causal.

An algorithm can appear consistent because it applies the same computational procedure repeatedly. Consistency is useful, but it is not the same as fairness. A system can reproduce the same mistake for every student, or apply a rule consistently to groups for whom the underlying data are less accurate. Removing visible emotion and discretion does not remove values from the grading design.

Responsibility can also become diluted. A human colleague who recommends five may be asked to defend the decision. An algorithm has no professional reputation to protect and cannot answer a student. If a teacher feels less blame belongs to the source, accepting the recommendation may become psychologically easier even though responsibility for the final grade still belongs to the school.

Automation bias is not blind obedience

The findings are better described as conditional deference than universal automation bias. Teachers had the work and an explicit checklist, and the algorithm label still moved their fairness judgement in one condition. Yet the absence of a significant lenient effect shows they did not simply prefer machine advice across the board.

Context may determine when deference appears. A strict score can be framed as rigour, protection of standards or resistance to grade inflation. The machine label may strengthen that frame by implying measurement rather than opinion. A generous score offers no equivalent signal of exacting standards, which may explain why its source mattered less. The experiment identifies the asymmetry but does not definitively establish this explanation.

LiveAIWire’s report on AI assistance hiding weaker human skills points to an adjacent challenge. AI-supported performance can look stronger than unaided ability, while an AI-supported judgement can look more objective than the reasoning underneath it. In both cases, institutions need separate evidence for what the human knows and what the system contributed.

Experienced professionals were not immune

The participants were teachers, not students or inexperienced crowd workers. That makes the study relevant to professional oversight claims. Simply placing a qualified person “in the loop” does not guarantee independent review if the system’s recommendation anchors the person before they evaluate the evidence.

Subgroup analysis suggested greater deference to harsh algorithmic advice among younger teachers, those with higher education and those who felt more confident with technology. Such comparisons should be interpreted cautiously because they divide the sample and can be sensitive to modelling choices. They nevertheless challenge the assumption that technical confidence automatically protects users from overreliance.

The better safeguard is procedural. Ask the human to make an initial assessment before revealing the automated score, then require a written reconciliation where the two differ materially. That design turns the model into a second opinion instead of an anchor. It also produces evidence about whether the tool adds value or merely shifts professional judgement.

Schools already face a wider assessment crisis

AI is entering assessment from both directions. Students use generative tools to produce work, while institutions use automated systems to detect, score and triage it. LiveAIWire’s investigation of the AI exam cheating crisis found that detection systems can miss machine-written answers and wrongly flag human work, especially writing by non-native English speakers.

Automated grading could intensify that problem if one uncertain system feeds another. A detector might estimate that an essay was AI-generated, a grader might lower its score, and a teacher might defer because both outputs appear technical. Each stage can be presented as advisory while the combined pipeline determines the student’s result.

Schools should therefore audit the complete decision path. Accuracy measured for one component does not establish fairness for the sequence. A review must ask which data enter, where uncertainty is displayed, when a person sees the recommendation, who may override it and whether errors fall disproportionately on particular students.

Adoption was common, enthusiasm was not

The survey portion of the research found that 48 per cent of participating teachers used AI weekly for lesson preparation, while 26 per cent said they never used it. Only 16 per cent would encourage colleagues to use AI. The profession represented in the experiment was therefore neither uniformly resistant nor uncritically enthusiastic.

That mixed picture makes the labelling effect more important. A person does not need to be an AI evangelist to grant a machine recommendation extra authority in a specific decision. Attitudes expressed in a survey and behaviour under time or workload pressure may diverge, especially when a score arrives in a polished interface.

AI tools can still help teachers by checking coverage of a rubric, identifying inconsistent feedback or producing a second reading for review. The value lies in surfacing information, not replacing judgement. A system should make disagreement easier by showing uncertainty and evidence, rather than presenting one number with the visual confidence of a settled fact.

The study does not prove real schools are doing this

The experiment used short vignettes and hypothetical grading decisions, not a deployed school platform affecting actual pupils. Participants knew the rubric, reviewed one task and did not receive the explanations, confidence scores or institutional incentives that accompany real systems. Behaviour could differ when teachers know a grade has consequences and may be appealed.

The sample came from Greece, so education culture, professional norms and technology policy may limit generalisation. The algorithm itself was only a label. That is a strength for identifying the influence of perceived source, but it says nothing about whether a particular grading model is accurate or how teachers respond after learning its track record.

The peer-reviewed article is indexed in PubMed and available in full through PubMed Central. The converging records support its publication and methods, but the result remains one controlled experiment. Replication across countries, subjects and live decisions is needed before estimating the scale of the problem.

Human review must begin before the machine verdict

Organisations often describe AI as safe because a professional retains final authority. The study shows why that claim is incomplete. A teacher can formally control the grade while still changing their perception of fairness in response to the system’s label. Oversight is a cognitive task, not a box in a workflow diagram.

Good review protects independent judgement. It hides the recommendation until an initial assessment is recorded, highlights large disagreements, gives the reviewer evidence rather than prestige cues and measures how often overrides improve outcomes. It also assigns responsibility clearly to the institution and the person signing off the decision.

This lesson fits broader evidence that human-AI teamwork does not automatically beat the best alternative. Combining two decision-makers can improve performance, but only when their roles are designed around complementary strengths. In grading, the human must be able to challenge the system before its score becomes the starting point.

The worrying result was not that teachers obeyed a robot. It was subtler: the word “algorithm” made the same harsh recommendation feel fairer. If schools want AI to support professional judgement, they must design for genuine disagreement. Otherwise the human in the loop may become the person who legitimises a number after the machine has already framed it.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.