Is AI Grading Fair? What the Evidence Says

We sell an AI grading tool, so treat what follows accordingly — but the case for reading it is that we are going to spend most of it on the evidence against fully automated marking, because that evidence is real and because pretending otherwise is how institutions get burned.
"Fair" is three questions, not one
When a teacher asks whether AI grading is fair, they are usually asking all three of these at once, and the answers are not the same:
- Consistency. Does the same piece of work get the same score on Monday and on Friday, and does the twentieth script get the same attention as the first?
- Construct validity. Is it scoring the thing the criterion describes, or something correlated with it — length, fluency, vocabulary, confident tone?
- Group fairness. Do two students of equal ability get equal scores when they differ in something irrelevant, such as their first language?
Automated marking tends to do well on the first, wobbles on the second, and is genuinely at risk on the third. Any vendor answering "is it fair?" with a single accuracy number is answering a question you did not ask.
How closely does AI agree with human markers?
The most useful single source here is a 2025 research synthesis, Agreement Between Large Language Models and Human Raters in Essay Scoring, which followed PRISMA 2020 and pulled together 65 published and unpublished studies from January 2022 to August 2025. Its headline is deliberately unglamorous: agreement between LLMs and human raters was generally moderate to good, with the reported indices mostly falling between 0.30 and 0.80 — across quadratic weighted kappa, Pearson correlation and Spearman's rho — and with substantial variability between studies.
Read that range honestly. The top of it is the level of agreement you would be pleased to see between two trained human markers. The bottom of it is close to useless. And the width of the range is the actual finding: performance is not a property of "AI grading", it is a property of a particular model, on a particular task, with a particular prompt, on a particular kind of writing. A result from a well-designed TOEFL study tells you very little about your third-year law essays.
The practical reading of that synthesis is not "AI marking works" or "AI marking doesn't". It is: whatever the vendor's benchmark says, the number that matters is the one you measure on your own assignments, with your own rubric.
Where the bias shows up
This is the part that should change how you deploy, not just whether you deploy.
A 2026 study by John Maurice Gayed took an open-weight model — Gemma-3-27B-it, LoRA-fine-tuned on 480 argumentative essays from two prompts — and evaluated it across eight unseen prompts on the TOEFL11 corpus: 12,100 essays from writers with 11 different first-language backgrounds. Overall performance looked respectable: 77.79% band agreement and a quadratic weighted kappa of 0.702, comfortably in the upper half of the range above.
The fairness analysis is the reason to care. The model showed a systematic, first-language-linked scoring offset: within every proficiency band, essays from European-language backgrounds were scored consistently higher than essays from East-Asian-language backgrounds — and the authors report that this was not explained by the composition of the training data.
Two things follow. First, an aggregate agreement figure can look fine while a group-linked offset sits underneath it, because the offset partly cancels out in the average. Second, the students most exposed are usually the ones with the least standing to complain — international students, second-language writers, anyone whose prose reads as unusual to a model trained mostly on one register of English.
AI does not grade like a human — even when it agrees
A 2026 paper with the admirably blunt title LLMs Do Not Grade Essays Like Humans (Mathew, Taher, Kundu and Barbosa) looked at behaviour rather than agreement, and found that models lean on different signals from human raters. In particular, they tend to award higher scores to short or underdeveloped essays, and lower scores to longer essays containing minor grammatical or spelling errors.
That combination is worth sitting with. A student who writes little and cleanly can be rewarded over a student who attempts something ambitious and makes surface slips doing it — which is close to the inverse of what most rubrics are trying to encourage. The same paper found that model scores are internally consistent with the feedback the model writes, and concluded that LLMs can still work as a supporting tool for essay scoring. Supporting is the operative word.
| Study | What it measured | What it found |
|---|---|---|
| Agreement Between LLMs and Human Raters in Essay Scoring: A Research Synthesis (2025) | 65 published and unpublished studies, Jan 2022 – Aug 2025, PRISMA 2020 | Agreement generally moderate to good; indices mostly 0.30–0.80, with substantial variability between studies |
| Investigating first-language bias in LLM-based automated essay scoring — Gayed (2026) | Fine-tuned Gemma-3-27B-it across 8 unseen prompts on TOEFL11 (12,100 essays, 11 L1 backgrounds) | 77.79% band agreement, QWK 0.702 — but a systematic L1-linked offset: European-language backgrounds scored above East-Asian-language backgrounds within every band |
| LLMs Do Not Grade Essays Like Humans — Mathew, Taher, Kundu & Barbosa (2026) | Which features drive LLM scores versus human scores | Higher scores for short or underdeveloped essays; lower scores for longer essays with minor grammar or spelling errors; scores internally consistent with the model's own feedback |
The comparison that actually matters
It is tempting to stop there and conclude that human marking is the safe option. That is not what the literature says either. Human marking is where most of the known assessment biases were documented in the first place, and it has failure modes an automated marker does not: fatigue across a large pile, drift in the standard between the first script and the fiftieth, and the ordinary human sensitivity to a name, a handwriting style, or the reputation of the student who wrote it.
So the honest comparison is not "AI versus a perfect marker". It is "AI-assisted marking versus the marking process you actually run in week 11, with the cohort you actually have". For a lot of departments that current process is one exhausted tutor at midnight, and its consistency has never been measured at all.
The reason to keep a human in the loop is not that humans are more accurate. It is that accountability has to sit somewhere a student can appeal to, and a model is not a somewhere.
What a fair-by-design setup looks like
None of the above argues against using AI in marking. It argues for a specific shape of use — one where the model is constrained, its reasoning is inspectable, and the decision stays with a person.
- The rubric comes from the instructor. A criteria-by-levels grid with written descriptors is the single most effective constraint on construct drift, because it tells the model what to look at instead of leaving it to form a general impression. In EduGears AI the rubric is yours either way: write it from scratch, start from the editable built-in default, or have the AI draft one and edit it — a draft is not a rubric until you save it, and nothing is attached to an activity until you attach it. The standard being applied to a student's work should be one a human chose and can point to.
- Score each criterion separately, with a justification. A single holistic number is unauditable. Per-criterion scores with a short written reason for each level let you see why a mark landed where it did, and make the length-and-fluency failure mode above visible instead of hidden.
- Hold the result as a draft. On the surfaces built for open-ended work — standalone assessments and project assignments — the AI score is a draft an instructor reviews before it can reach the gradebook, and any change is recorded as a teacher edit. On assessments there is no auto-post option at all.
- Apply the same rubric to everyone. Consistency is the one dimension automation is genuinely good at. It is worth having, and it is worth not throwing away by running some students through the tool and others by hand.
- Say so, publicly. Tell students AI assists the marking, that a human approves every score, and how to request a re-mark. Most fairness complaints are really transparency complaints arriving late.
Instructor-in-the-loop is not a hedge or a legal disclaimer. Given a body of evidence that says agreement varies widely by context and that group-linked offsets can hide under a decent average, an approval step is the only mechanism that catches the case your benchmark did not cover.
How to check your own courses
You do not need a research team for this. A single afternoon per course gets you most of the value.
- Take a set you have already marked by hand — 30 scripts is enough to see a pattern — and run it through the tool without looking at your own scores first.
- Plot the difference, not the average. A mean difference near zero can hide a strong pattern in both directions.
- Sort the disagreements by student group where you legitimately hold that data — first language, entry pathway, mode of study. If one group is consistently on one side of the line, you have found the thing the research warns about, and you have found it before a student did.
- Look at the extremes. The shortest scripts and the longest ones are where the documented failure modes live.
- Read three justifications in full. If the reasoning does not match the criterion, the rubric wording is the problem more often than the model is.
- Repeat when the model changes. An underlying model update can shift behaviour without anything in your course changing, so re-run the check at least once a year.
One thing we are not going to do here is give you an accuracy figure for our own grading. We have not run a published fairness audit of it, and a number without a methodology behind it would be worth less than the paragraph above telling you how to measure it yourself.
Rubric-based AI marking with per-criterion justifications and a human approval step, inside Moodle, Canvas, Blackboard or Brightspace.
Get started free →Frequently asked questions
Is AI grading accurate enough to replace a human marker?
On the published evidence, no — not for open-ended work. A 2025 synthesis of 65 studies found agreement between LLMs and human raters was generally moderate to good, with indices mostly between 0.30 and 0.80 and substantial variability between studies. The top of that range is comparable to two trained human markers; the bottom of it is not usable. Because the result depends so heavily on the model, the task and the kind of writing, a vendor benchmark tells you little about your own assignments. Use it to draft and to enforce consistency, and keep a person making the decision.
Is AI grading biased against second-language writers?
There is direct evidence that it can be. A 2026 study fine-tuned an open-weight model and evaluated it on 12,100 TOEFL essays from writers with 11 different first-language backgrounds. Overall agreement looked healthy at 77.79% band agreement and a quadratic weighted kappa of 0.702, but within every proficiency band, essays from European-language backgrounds scored consistently higher than those from East-Asian-language backgrounds, and the authors report this was not attributable to the training data composition. The lesson is that an aggregate accuracy figure can look fine while a group-linked offset sits underneath it, so check the disagreements by group rather than trusting the mean.
Do AI markers just reward long, well-spelled essays?
The documented pattern is stranger than that, and worth knowing. A 2026 paper found LLM scorers tend to award higher scores to short or underdeveloped essays, and lower scores to longer essays that contain minor grammatical or spelling errors — which can penalise a student for attempting something ambitious. The same work found model scores are internally consistent with the feedback the model writes. That consistency is exactly why per-criterion justifications are useful: if the reasoning does not match your criterion, you can see it in the text rather than guessing from a number.
Isn't human marking biased too?
Yes, and that is the fair comparison. Most of the known assessment biases were documented in human marking long before automated scoring existed, and human marking has failure modes a machine does not — fatigue over a large pile, drift in the standard between the first script and the fiftieth, and sensitivity to a name or a handwriting style. The reason to keep a human in the loop is not that humans mark more accurately. It is that a student needs somewhere to appeal, and accountability has to rest with a person who can explain and change a decision.
What makes an AI grading setup defensible to a moderator?
Four things, and all of them are about the record rather than the model. A rubric a named human wrote, with criteria and level descriptors you can show. A score per criterion rather than one holistic number, each with a written justification. A draft-and-approve step, so the gradebook reflects a decision an instructor made and any change is recorded as a teacher edit. And a stated policy students can read before they submit. If a mark is challenged, those four give you an audit trail; a single AI-generated number gives you nothing to point at.
How often should we re-check a marking tool?
At least once a year, and always after the underlying model changes — an AI provider can update a model without anything in your course changing, and behaviour can shift with it. The check itself is cheap: take about 30 scripts you have already marked, run them through without looking at your own scores, plot the differences rather than the average, and sort the disagreements by student group where you legitimately hold that data. If one group sits consistently on one side of the line, you have found the problem before a student did.
Related reading: can Moodle grade essays with AI, and why rubric-based grading is the control surface.
Try EduGears AI Free
Setup in 3 minutes via LTI 1.3. No credit card required. All 26 AI tools included on the free tier.
Get Started Free →