AI Grading

AI Rubric Generation, Done Right

EduGears AI Team··13 min read
Flat brand illustration of rubric-based AI grading — an assignment brief and deliverable checklist alongside a criteria-by-levels rubric grid marked as an AI draft awaiting review, with edit and approval badges, flowing into an LMS gradebook.

Every conversation about AI grading eventually arrives at the same question: how do I know the score is right? The answer is not a better model. It is a better rubric. A rubric is the artefact that turns a model's opinion into a score you can explain to a student, a colleague, or an appeals committee — and it is the one part of the workflow that stays entirely under your control.

For most of this year the version of this article you would have found here was titled Why We Don’t Do It. We had looked at AI rubric generation as the category is usually sold — a model writes the criteria, the weightings and the descriptors, and the teacher inherits a standard nobody in the room actually chose — and decided we would rather ship nothing than ship that. What changed is not the argument. It is that we found a shape that keeps the argument intact: AI drafts, you own the standard. Draft with AI is now live in the Rubrics tool, and everything it produces is a draft that does not exist as a rubric until a teacher edits it and saves it.

This guide covers what AI rubric generation usually does and where it goes wrong, how Draft with AI works and what it deliberately refuses to do, how to write criteria a model applies consistently, and what happens once the rubric is set — per-criterion scoring, instructor review, and grade passback into Moodle, Canvas, Blackboard, or Brightspace. If you are still choosing a tool, start with our comparison of the best AI grading software for higher education; if you want the LMS plumbing in detail, see AI grading in Moodle, Canvas & Blackboard.

Why Is the Rubric the Control Surface for AI Grading?

The rubric is the control surface for AI grading because it is the only thing that constrains what the model is allowed to evaluate and how much it is allowed to award. A rubric with explicit criteria, named performance levels, and point values converts an open-ended judgement into a bounded decision: pick one level per criterion, from a fixed set, worth a fixed number of points. Without that structure you get confident-sounding scores that drift between submissions and cannot be defended when a student asks why they lost four marks.

Three specific things a rubric does that no amount of prompt engineering replaces:

  • It sets the ceiling. If a criterion tops out at four points, four points is the most that criterion can be worth — the model cannot invent a fifth level or award bonus marks for something it liked.
  • It sets the vocabulary. Criteria names and level descriptors are the language the feedback comes back in, so students and instructors argue about the same thing rather than about “quality”.
  • It makes the score auditable. A per-criterion breakdown shows exactly where marks were lost, which means a disagreement is a conversation about one criterion rather than about a single opaque number.

This is also why the rubric is the right place to spend your time. An hour refining criteria pays back across every submission in the cohort, every subsequent semester the rubric is reused, and every appeal you do not have to litigate from scratch.

The useful mental model: the rubric is the specification and the AI is the implementation. A vague specification produces an unpredictable implementation — and the fix is always to tighten the specification, never to argue with the model.

What AI Rubric Generation Actually Does

AI rubric generation, as the category is usually sold, is the use of a language model to draft a rubric — criteria, performance levels, descriptors, and point values — from the context of an assignment, so that an instructor edits a draft instead of starting from an empty grid. It is worth understanding properly, because the pitch is genuinely appealing and the output is a proposal rather than a finished instrument. Generation solves the blank-page problem. It does not solve the judgement problem, and the judgement problem is the only one that matters when a student challenges a mark.

What the rubric has to reflect, whoever writes it

This is the part generation vendors are right about, and it is worth keeping even if you write every criterion by hand. A generic essay rubric can be pulled off a shelf in seconds and is almost always wrong for the assignment in front of you. What makes a rubric fit is:

  • The assignment brief itself — the actual task text, not a topic label. “Compare two treatment protocols and justify a recommendation” produces different criteria than “write about treatment protocols”.
  • The deliverables — what students physically hand in. A rubric for a report, a poster, and a codebase should not share a mechanics criterion.
  • Learning outcomes — the course or module outcomes the assessment is meant to evidence, so the criteria map to something a programme review can trace.
  • Level and discipline — first-year undergraduate versus postgraduate changes what “critical analysis” means in a descriptor.
  • Shape constraints — how many criteria, how many performance levels, and the point scale, so the draft fits the grading scheme you already use.

What a generated draft is good at, and what it is not

A generated draft reliably gets you a plausible spread of criteria, a consistent number of levels with parallel phrasing, and descriptors that are at least grammatically distinguishable. What it routinely gets wrong is the two things that decide whether a rubric is any good. Weighting: models tend to distribute points evenly across criteria, which almost never reflects what the assignment is actually assessing. Specificity: they produce descriptors like “demonstrates good understanding” that read fine and cannot be applied consistently by anyone, human or otherwise.

So the honest accounting of a generated rubric is: it saves you the typing, and leaves you the thinking. And the thinking is most of the hour. That accounting is not an argument against drafting — it is an argument against calling a draft a rubric, and it is the seam the next section is built along.

How Draft with AI Works in EduGears AI

Worth being precise here, because this is a place vendors tend to overclaim. EduGears AI drafts rubrics; it does not decide your standard. The Draft with AI button sits next to New rubric in the Rubrics tool, and what it returns is a filled-in criteria-by-levels grid sitting in the ordinary rubric editor, unsaved. No rubric exists yet, nothing is attached to an activity, and no student is graded against anything until an instructor reads it and presses Create rubric themselves — the same button a hand-built rubric goes through, with the same validation.

That constraint is the whole design, and it is the same argument the previous section makes, applied to ourselves. The standard being applied to a student’s work has to be one a human chose and can point to. The moment a model proposes the criteria and applies them with nobody in between, the only thing anchoring the score to your course is the model’s own opinion of what your assignment was about — and there is nobody to appeal to. What we were never willing to ship is auto-attachment: a rubric arriving on an activity without a teacher having read it. Drafting saves the typing; the review step keeps the thinking exactly where it was.

What the dialog asks for

Draft with AI opens a short dialog rather than firing on a single click, because a draft is only as good as the context it was given:

  • Assignment title and what students are being asked to do — a title on its own works; a brief works better. The more specific you are about the task, the evidence students must use, and what a strong response looks like, the less of the draft you will rewrite.
  • Category — quiz, essay, student assessment, or project, the same four categories rubrics are already filed under, which is also what decides where the finished rubric appears in the activity pickers.
  • Criteria and levels — how many rows and columns to propose. Four of each is the default; you can ask for between two and twelve criteria and between two and eight levels, so the draft arrives in the shape of the grading scheme you already use.
  • Use my course source materialoff by default. Left off, the draft is written from your title and brief alone. Turned on, it is also grounded in the files you have uploaded to this course, so criteria borrow the vocabulary of your material instead of generic language. If the course has no uploaded files, the draft is written ungrounded and says so.

Drafting is instructor-only, and that is enforced on the server rather than by hiding a button. While the draft is being written the dialog announces progress, and a failure leaves your inputs in place so you can retry without retyping them.

What comes back, and the banner that stays up

The draft lands in the normal rubric editor under a standing banner reading AI draft — review every criterion before you save. It is not a dismissible toast: it stays for as long as the draft is unsaved, and clears on the first save. Everything in the grid is editable in place — rename criteria, rewrite descriptors, change the points, add or remove rows and columns. Read the weightings especially, because an evenly distributed point spread is the most predictable thing a model does and it is rarely what your assignment actually assesses. The draft is not a rubric until you save it, and saving is what makes it yours.

The AI Grading tool’s built-in default rubric has not gone anywhere either, and it is still not model output — content accuracy and relevance (30%), completeness (25%), critical thinking and analysis (20%), clarity and organisation (15%), and grammar, spelling and formatting (10%) — which you can apply as-is to general written work or edit before grading. It is a fixed template we wrote, identical for everyone, and visible in full before you use it — you can read it, disagree with it, and change it, and it says the same thing on Tuesday as it did on Monday. So there are three ways to start a rubric now: an empty grid, a template you can read end to end, or an AI draft you review. All three land in the same editor, and all three are saved by you.

Wherever you start, the durable version lives in the Rubrics tool as a structured criteria-by-levels grid: each criterion has a name and description, each level has a name, a descriptor, and a point value, and the maximum score is derived by summing the highest level of every criterion. A criterion needs at least two levels to be meaningful, so that is enforced — on an AI draft exactly as on a hand-built rubric. A rubric is reusable — build it once, attach it to as many activities as you like — and a visibility toggle controls whether learners see it on the activity page. Rubrics are categorised by what they grade: quiz, essay, student assessment, or project. Attaching one is always a separate, deliberate step taken from the activity itself: nothing an AI drafts reaches a course on its own.

One consequence of reuse worth knowing before you edit: a rubric is attached to activities by reference, so changing a criterion changes grading everywhere that rubric is used. If you want to revise mid-semester without disturbing work already graded, copy the rubric and attach the copy.

How to Write Rubric Criteria a Model Can Apply

Write rubric criteria a model can apply by making each criterion observable, independent, and bounded — the same properties that make a rubric usable by a new teaching assistant. If two markers with your rubric and no other context would land on the same level, a model will too. The practical rules:

  1. Name the evidence, not the virtue. Replace “Critical thinking” with “Evaluates at least two competing explanations and states which is better supported”. The first is a compliment; the second is something you can point at in the text.
  2. Make levels differ on one dimension. A descriptor set that changes quality, quantity, and accuracy between adjacent levels forces an arbitrary choice. Vary one axis — usually completeness or depth — and hold the rest constant.
  3. Anchor the extremes first. Write the top and bottom descriptors, then fill the middle. Middles written first tend to be hedged, and everything above and below inherits the hedge.
  4. Keep criteria independent. If weak structure drags down three criteria, one weak submission loses marks three times. Overlapping criteria are the single most common cause of scores that feel unfairly harsh.
  5. Weight deliberately. Points are the honest statement of what the assignment assesses. If mechanics carries the same weight as argument, the rubric says mechanics matters as much as argument — decide whether you mean that.
  6. Say what a zero looks like. The lowest level should describe work that was attempted but missed, not absence. Missing work is a submission problem, not a rubric level.

Refining a rubric after the first run

The fastest way to improve a rubric is to grade a handful of submissions and read the AI's per-criterion justifications rather than its scores. Justifications expose ambiguity immediately: if the rationale for a criterion keeps restating the descriptor instead of citing something in the submission, the criterion is not observable. If the model keeps choosing the same middle level across very different submissions, the levels do not discriminate. If it awards a high level for something you would have marked down, the descriptor is describing a different thing than you meant. Fix the criterion, regrade the sample, and only then run the cohort.

A worked example

Weak criterionWhy it failsRewritten
Quality of researchNot observable; two markers will disagree about what countsCites at least five peer-reviewed sources published in the last ten years, each used to support a specific claim
Good structure and flowTwo dimensions in one criterion; overlaps with argumentEach section opens with a claim and closes by linking to the next section
Demonstrates understanding of the topicRestates the assignment; every level would read the sameDefines the three core terms accurately and applies each to the case study without contradiction

How AI Scores Against Your Rubric, Criterion by Criterion

Once a rubric is attached to an activity, AI scoring in EduGears AI works one criterion at a time rather than producing a single overall mark. For each criterion the model selects one of the performance levels you defined, assigns the points attached to that level, and writes a short justification for the choice. It then adds an overall comment, and on project and assessment submissions an explicit list of strengths and improvements.

Three guardrails matter more than the scoring itself:

  • Scores are clamped to the criterion maximum. The point ceiling you set in the rubric is the ceiling that gets applied, so a model cannot inflate a criterion beyond its top level.
  • The breakdown always mirrors the rubric. If the model skips a criterion, it comes back at zero points marked as not graded by AI rather than quietly vanishing — so a missing judgement is visible instead of invisible.
  • The model cannot introduce criteria. Scoring is restricted to the criteria in your rubric; anything the model has an opinion about that you did not ask for has nowhere to go.

Submissions are read directly rather than summarised first. PDFs, plain text, and images are handled natively, and DOCX and PPTX uploads are text-extracted so their contents are graded rather than their filenames. That is what makes rubric grading workable for photographed lab notebooks, scanned problem sets, and slide decks, not just typed essays.

Two operational details worth knowing. A failed AI call falls back to manual review rather than leaving a submission in limbo, so a provider outage costs you a grading pass and nothing else. And if you connect your own provider account with the bring-your-own-key option, grading runs on your key at your provider's rate, which is what most institutions with an existing Azure OpenAI or Anthropic contract end up doing.

What You Review Before a Grade Posts

In EduGears AI, an AI score is held as a draft and does not reach the student gradebook until an instructor acts. AI Grading and Student Assessment submissions are always held for review — there is no automatic mode to turn on. Project and capstone assignments default to draft mode, where you accept the AI's scores as they stand or override any individual criterion, and can optionally be switched to automatic for low-stakes work where you have decided review is not worth the time.

What to actually look at during review, in the order that catches the most problems:

  1. Any criterion marked as not graded by AI. These are the submissions where the model had no confident judgement — usually a signal about the criterion, not the student.
  2. Justifications that do not quote the work. A rationale that paraphrases your descriptor rather than pointing at the submission is a level chosen on vibes.
  3. Boundary cases. Submissions sitting between two levels on a heavily weighted criterion are where a one-level change moves the grade most, and where your judgement adds the most value.
  4. The extremes. Read the top and bottom of the cohort in full. If both look right, the middle usually is.

Overrides work the same way manual grading does, and they are bounded by the same rubric — a teacher-entered score is clamped to the criterion maximum too, so the rubric stays the single source of truth for what an assignment is worth.

How the Approved Score Reaches Your Gradebook

Once you approve a score, it posts to your LMS gradebook over LTI Assignment and Grade Services (AGS), the part of the LTI 1.3 Advantage specification that lets an external tool write results back into a course. There is no CSV export, no copy-paste, and no nightly sync job: the tool posts the score and comment to the LMS, the LMS validates the credential, and the gradebook row updates. This works the same way across Moodle, Canvas, Blackboard, and Brightspace, because the standard is the same everywhere.

One behaviour worth knowing about project assignments: instead of posting to whatever line item the learner happened to launch through, a project score posts to a dedicated gradebook column created for that project. That means a project grade can never overwrite a course-completion column — a genuinely unpleasant failure mode when a tool gets it wrong.

Draft a rubric with AI, edit it until the standard is yours, grade a cohort against it — free, no credit card.

Get Started Free →

Rubrics for Project and Reflection Assignments

Open-ended work is where rubrics earn their keep, and also where most rubrics are worst. A project or a reflection has no answer key, so the rubric is doing all of the work of defining what good looks like — and a rubric written for an essay usually transfers badly.

Project rubrics: grade the artefact, not the write-up

The failure mode in project grading is a rubric that accidentally scores the description of the work rather than the work. If students submit files plus a written write-up, at least half the criteria should be about the submitted artefacts — does the prototype do what it claims, is the dataset documented, does the design meet the brief's constraints. In EduGears AI a project assignment carries a brief, a deliverable checklist, and a rubric, and all three feed the grading pass, so a criterion can legitimately reference something the checklist required. A rubric is required before a project can be activated at all, which is a deliberate constraint: you cannot ship an open-ended assignment without first saying how it will be judged.

Publish the rubric to students before they start. On open-ended work this measurably changes what you get back, because most of what separates a strong project from a weak one is knowing what was being asked for. Our guide to AI-graded project assignments covers the full workflow.

Reflection rubrics: assess the thinking, never the feeling

Reflective writing is the hardest thing on this list to rubric well, because the obvious criteria — depth, honesty, insight — are exactly the ones that are not observable, and because grading a student's stated feelings is both unfair and unmeasurable. The workable approach is to assess the structure of the reflection rather than its content: whether a specific incident is described concretely, whether the student connects it to a named concept from the course, whether they identify what they would do differently and why, and whether the reflection engages with something that did not go well rather than only successes.

Two practical constraints. Keep reflection rubrics short — three or four criteria — because long rubrics push students into checklist-completion and out of actual reflection. And use fewer levels, often three rather than five, since finer gradations on subjective criteria produce false precision that neither you nor a model can apply consistently.

Frequently Asked Questions

What is AI rubric generation?

AI rubric generation is the use of a language model to draft a grading rubric — criteria, performance levels, descriptors, and point values — from the context of an assignment, so an instructor edits a draft rather than starting from an empty grid. The output should always be treated as a proposal rather than an instrument: models typically distribute points evenly across criteria, which rarely reflects what an assignment actually assesses, and produce descriptors that read well but cannot be applied consistently by anyone. In practice that means generation saves you the typing and leaves you the thinking, and the thinking is most of the work. EduGears AI offers it in the only form we think is defensible: Draft with AI proposes the grid, and it stays an unsaved draft until an instructor edits it and saves it, so the standard reaching a student is one a human chose.

Can EduGears AI generate a rubric from my assignment?

Yes — as a draft you own, and never automatically. <strong>Draft with AI</strong> in the Rubrics tool takes an assignment title, a brief, the rubric category, and how many criteria and levels you want, and returns a filled-in criteria-by-levels grid in the ordinary editor. Grounding the draft in your uploaded course material is an explicit toggle that is off by default. What comes back is an unsaved draft under a standing “AI draft — review every criterion before you save” banner: no rubric exists, nothing is attached to any activity, and no student is graded until you edit it and press Create rubric yourself. Drafting is instructor-only. The built-in default rubric in the AI Grading tool is still there too — content accuracy and relevance, completeness, critical thinking and analysis, clarity and organisation, and grammar, spelling and formatting at 30/25/20/15/10, a fixed template we wrote, identical for every user. Once your rubric is saved, AI grading scores every submission against it criterion by criterion, with a justification for each level it selects.

Can the AI give a score higher than my rubric allows?

No. Scores are clamped to each criterion's maximum, so the point ceiling defined by your top performance level is the most that criterion can award. The AI also cannot score against criteria you did not define — the breakdown always mirrors the shape of your rubric. If the model fails to produce a judgement for a criterion, that criterion is returned at zero points and explicitly marked as not graded by AI, so a missing judgement is visible to you during review rather than silently absent. Teacher overrides are bounded by the same maximums, which keeps the rubric the single source of truth for what an assignment is worth.

Do I have to review every AI grade before students see it?

For AI Grading and Student Assessment submissions in EduGears AI, yes — those are always held as a draft for instructor review, and there is no automatic mode to enable. Project and capstone assignments default to draft mode, where you accept the AI's scores or override any criterion before posting, and can optionally be switched to automatic for low-stakes work. Nothing reaches the LMS gradebook until a human action posts it. In practice the efficient review pattern is to check criteria marked as not graded by AI, justifications that do not cite the submission, boundary cases on heavily weighted criteria, and the top and bottom of the cohort in full.

How do rubric grades get into the Moodle or Canvas gradebook?

Approved rubric scores post over LTI Assignment and Grade Services (AGS), part of the LTI 1.3 Advantage specification, which lets an external tool write results directly into an LMS course. There is no CSV export, no copy-paste, and no nightly sync job — the tool posts the score and comment, the LMS validates the credential, and the gradebook updates. Because AGS is a standard, the behaviour is the same across Moodle, Canvas, Blackboard, and Brightspace. Project assignments in EduGears AI post to a dedicated gradebook column created for that project, so a project score can never overwrite the course-completion column a learner launched through.

How many criteria and levels should a rubric have?

For most written assignments, four to six criteria and four performance levels is a workable default: enough criteria to give a useful breakdown, few enough that each one carries meaningful weight. Reflective assignments work better with three or four criteria and often three levels, because finer gradations on subjective criteria produce false precision that neither an instructor nor a model can apply consistently. A criterion needs at least two levels to discriminate at all. The more important constraint is independence — if one weakness in a submission costs marks under three different criteria, the rubric will produce scores that feel unfairly harsh regardless of how many criteria it has.

Rubrics are one of 26 AI tools that install together in about three minutes over LTI 1.3, so rubric-based grading — against your rubric — sits alongside question generation, quizzes, project assignments, and the rest of the toolkit with nothing extra to procure. To compare approaches before you commit, read the best AI grading software for higher education, or follow the setup guide for your LMS.

Try EduGears AI Free

Setup in 3 minutes via LTI 1.3. No credit card required. All 26 AI tools included on the free tier.

Get Started Free →

Related posts