Every question is graded against a rubric, stored in the question's scoring object. The rubric describes what earns points, and Examplary's auto-grading AI applies it to each answer: it picks the criteria or levels that fit, awards points, and explains its reasoning per criterion.
There are four rubric types, set with rubricType:
| Rubric type | How it's scored | Maximum points | Graded by |
|---|---|---|---|
simple | Each criterion that applies adds its points. | The sum of all criteria | AI |
analytical | One level is picked per criterion; the levels' points add up. | The sum of each criterion's top level | AI |
holistic | One criterion is picked for the answer as a whole. | The highest criterion | AI |
exact-values | The answer must match one of the accepted values exactly. | The highest criterion | Automatic |
The full structure is available as a JSON Schema, which you can use to validate rubrics or generate types in your own system. It's also published on npm as part of @examplary/schemas.
The scoring object
{
"rubricType": "analytical",
"criteria": [ … ],
"guidance": "Accept answers that mention the number 42 without the book title.",
"modelAnswer": "42, the answer from The Hitchhiker's Guide to the Galaxy."
}| Field | Type | Required | Description |
|---|---|---|---|
rubricType | string | No | simple (default), analytical, holistic or exact-values. |
criteria | array | Yes | The rubric's criteria. Their shape depends on the rubric type, see below. |
guidance | string | No | Extra grading instructions. The AI takes these into account for every answer. |
modelAnswer | string | No | An example of a full-marks answer. The AI uses it as a reference. |
exactValuesCaseInsensitive | boolean | No | Only for exact-values: ignore capitalisation and whitespace when matching. |
excludeFromGrading | boolean | No | Leaves the question out of the session's total points. It's still graded. |
templateId | string | No | The workspace rubric template this rubric was created from, if any. |
Each criterion, and each level of an analytical criterion, has these fields:
| Field | Type | Required | Description |
|---|---|---|---|
id | string | Yes | A unique ID. You'll find it again in the grading results, so pick something you can recognise. |
title | string | analytical, holistic | A short name. Used for analytical criteria and levels, and holistic criteria. |
description | string | simple, holistic, exact-values | What an answer needs to earn the points. For exact-values, the accepted value itself. |
points | number | Yes | The points awarded when the criterion or level applies. |
minPoints | number | No | Optional lower bound, when the criterion or level allows a range of points. |
maxPoints | number | No | Optional upper bound for that range. |
levels | array | No | Only for analytical criteria: the levels to choose from, each with the fields above. |
Examplary tidies up rubrics when you save a question: missing ids are
generated, fields that don't apply to the rubric type are dropped (such as
levels outside analytical rubrics, or title on simple criteria), and
analytical levels and holistic criteria are sorted by points. Read the
question back after saving if you need the final shape.
Simple
A checklist: each criterion describes something the answer should contain, and every criterion that applies adds its points. Use it for answers that need to cover several separate points.
{
"rubricType": "simple",
"criteria": [
{
"id": "mentions-42",
"description": "Mentions the number 42.",
"points": 1
},
{
"id": "names-book",
"description": "Names The Hitchhiker's Guide to the Galaxy.",
"points": 1
},
{
"id": "explains-joke",
"description": "Explains why the answer is a joke.",
"points": 2
}
]
}The maximum is the sum of all criteria, 4 points here. In the results, every criterion gets its own entry with the points the AI awarded for it:
{
"gradingStatus": "auto-graded",
"gradedBy": "system-ai",
"pointsAwarded": 2,
"feedback": "<p>You found the answer, but didn't say where it comes from.</p>",
"scoringCriteria": [
{
"id": "mentions-42",
"pointsAwarded": 1,
"aiPointsAwarded": 1,
"confidenceScore": 95,
"reasoning": "States 42 directly."
},
{
"id": "names-book",
"pointsAwarded": 0,
"aiPointsAwarded": 0,
"confidenceScore": 90,
"reasoning": "No reference to the book."
},
{
"id": "explains-joke",
"pointsAwarded": 1,
"aiPointsAwarded": 1,
"confidenceScore": 60,
"reasoning": "Hints at the joke but doesn't explain it."
}
]
}Analytical
A grid: each criterion is a dimension of the answer, such as correctness or argumentation, with levels from weak to strong. The AI picks exactly one level per criterion, and the levels' points add up to the score. Analytical criteria only have a title and levels; the points live on the levels.
{
"rubricType": "analytical",
"criteria": [
{
"id": "correctness",
"title": "Correctness",
"levels": [
{
"id": "correctness-0",
"title": "Incorrect",
"description": "Doesn't mention 42.",
"points": 0
},
{
"id": "correctness-1",
"title": "Partially correct",
"description": "Mentions 42, but not the book.",
"points": 1
},
{
"id": "correctness-2",
"title": "Correct",
"description": "Mentions 42 and the book.",
"points": 2
}
]
},
{
"id": "clarity",
"title": "Clarity",
"levels": [
{ "id": "clarity-0", "title": "Unclear", "points": 0 },
{ "id": "clarity-1", "title": "Clear", "points": 1 }
]
}
]
}The maximum is the sum of each criterion's highest level, 3 points here. In the results, each criterion's entry also says which level was picked, in selectedLevelId:
{
"gradingStatus": "auto-graded",
"gradedBy": "system-ai",
"pointsAwarded": 2,
"feedback": "<p>Correct number! Mention where it comes from for full marks.</p>",
"scoringCriteria": [
{
"id": "correctness",
"selectedLevelId": "correctness-1",
"aiSelectedLevel": "correctness-1",
"pointsAwarded": 1,
"aiPointsAwarded": 1,
"confidenceScore": 85,
"reasoning": "Mentions 42, but not the book."
},
{
"id": "clarity",
"selectedLevelId": "clarity-1",
"aiSelectedLevel": "clarity-1",
"pointsAwarded": 1,
"aiPointsAwarded": 1,
"confidenceScore": 90,
"reasoning": "A short, direct answer."
}
]
}Holistic
A single scale for the answer as a whole: each criterion describes one overall quality, and the AI picks the one that fits best. Use it when an answer is best judged in one go, like a short essay.
{
"rubricType": "holistic",
"criteria": [
{
"id": "insufficient",
"title": "Insufficient",
"description": "Misses the point of the question.",
"points": 0
},
{
"id": "adequate",
"title": "Adequate",
"description": "Gives the right answer without context.",
"points": 2
},
{
"id": "excellent",
"title": "Excellent",
"description": "Gives the answer and explains its origin.",
"points": 4
}
]
}The maximum is the highest criterion, 4 points here. In the results, every criterion has an entry, but only the picked one has points; the others are 0:
{
"gradingStatus": "auto-graded",
"gradedBy": "system-ai",
"pointsAwarded": 2,
"scoringCriteria": [
{
"id": "insufficient",
"pointsAwarded": 0,
"aiPointsAwarded": 0,
"confidenceScore": 90,
"reasoning": "…"
},
{
"id": "adequate",
"pointsAwarded": 2,
"aiPointsAwarded": 2,
"confidenceScore": 80,
"reasoning": "Right answer, no context."
},
{
"id": "excellent",
"pointsAwarded": 0,
"aiPointsAwarded": 0,
"confidenceScore": 85,
"reasoning": "…"
}
]
}To find the picked criterion, look for the entry with points. If a teacher changes the grade in Examplary, the results instead hold just the picked criterion, with selectedLevelId set to its own id.
Exact values
For answers with a known correct value, like fill-in-the-blank questions. Each criterion's description is an accepted value, with the points it's worth. There's no AI involved: the answer is graded straight away when it's saved.
{
"rubricType": "exact-values",
"exactValuesCaseInsensitive": true,
"criteria": [
{ "id": "full", "description": "42", "points": 2 },
{ "id": "written-out", "description": "forty-two", "points": 1 }
]
}The first matching criterion wins, and the maximum is the highest criterion, 2 points here. Answers are compared after trimming surrounding whitespace. With exactValuesCaseInsensitive, capitalisation and all whitespace are ignored too, so "New York" matches "newyork". See exact-value matching for the details.
The results are final straight away, and hold the matching criterion only:
{
"gradingStatus": "graded",
"gradedBy": "system-exact-values",
"pointsAwarded": 2,
"scoringCriteria": [
{ "id": "full", "selectedLevelId": "full", "pointsAwarded": 2 }
]
}An answer that matches nothing gets "pointsAwarded": 0 and an empty scoringCriteria.
Reading the results
Grading results are part of each answer in the assessment session, which you get from the sessions endpoints or the exam.session.grading.completed webhook. See step 4 of the developer guide for how to fetch them.
| Field | Description |
|---|---|
gradingStatus | auto-grading-pending, auto-graded (an AI suggestion, not yet reviewed), auto-grading-failed, graded (final), or not-graded. |
gradedBy | system-ai, system-exact-values, system-response-processing (graded by the question type), or the ID of the teacher who graded it. |
pointsAwarded | The answer's points, rounded to one decimal. |
pointsToEarn | The question's maximum points. |
feedback | Feedback for the student, as HTML (see below). Not always present. |
scoringCriteria | The score per criterion, matched to your rubric by id. See below. |
gradedAt | When the answer was graded. |
Each entry in scoringCriteria has:
| Field | Description |
|---|---|
id | The criterion's id from the rubric. |
pointsAwarded | The points for this criterion. |
selectedLevelId | The picked level, for analytical criteria (and the matched value, for exact values). |
reasoning | The AI's explanation of the score, as HTML (see below). Useful to show to teachers next to the grade. |
confidenceScore | How sure the AI is, from 0 to 100. |
aiPointsAwarded | What the AI originally awarded. Stays the same when a teacher changes pointsAwarded. |
aiSelectedLevel | The level the AI originally picked, for analytical criteria. |
Both feedback and reasoning can contain HTML formatting, plus two custom tags: <inline-math> for formulas, and <content-reference> to quote part of the student's answer. Its quote attribute holds the exact text it refers to:
The student mentioned <content-reference quote="42">"42"</content-reference> but
did not name the book.When showing these in your own interface, you can highlight the quoted text in the answer, or just render the tag's content as plain text.
The AI's grades are suggestions, so they come back as auto-graded and count towards the session's points but not its final grade. A teacher can review them in Examplary, or you can finalise them yourself: accept a suggestion as it is, accept all of a session's suggestions, or save your own grade. Either way the answer becomes graded, and a late AI result never overwrites it.
Clear, specific criterion and level descriptions make the biggest difference.
Add guidance for anything that applies across criteria, and a modelAnswer
to show what full marks looks like. To teach the AI your own grading style,
add previously graded answers as
examples.