Examplary
  • Start for free
    Developer docs

    Rubrics & rubric types

    How rubrics are structured, how each rubric type is scored, and how grading results come back from the API.

    Every question is graded against a rubric, stored in the question's scoring object. The rubric describes what earns points, and Examplary's auto-grading AI applies it to each answer: it picks the criteria or levels that fit, awards points, and explains its reasoning per criterion.

    There are four rubric types, set with rubricType:

    Rubric typeHow it's scoredMaximum pointsGraded by
    simpleEach criterion that applies adds its points.The sum of all criteriaAI
    analyticalOne level is picked per criterion; the levels' points add up.The sum of each criterion's top levelAI
    holisticOne criterion is picked for the answer as a whole.The highest criterionAI
    exact-valuesThe answer must match one of the accepted values exactly.The highest criterionAutomatic

    The full structure is available as a JSON Schema, which you can use to validate rubrics or generate types in your own system. It's also published on npm as part of @examplary/schemas.

    The scoring object

    scoring
    {
      "rubricType": "analytical",
      "criteria": [ … ],
      "guidance": "Accept answers that mention the number 42 without the book title.",
      "modelAnswer": "42, the answer from The Hitchhiker's Guide to the Galaxy."
    }
    FieldTypeRequiredDescription
    rubricTypestringNosimple (default), analytical, holistic or exact-values.
    criteriaarrayYesThe rubric's criteria. Their shape depends on the rubric type, see below.
    guidancestringNoExtra grading instructions. The AI takes these into account for every answer.
    modelAnswerstringNoAn example of a full-marks answer. The AI uses it as a reference.
    exactValuesCaseInsensitivebooleanNoOnly for exact-values: ignore capitalisation and whitespace when matching.
    excludeFromGradingbooleanNoLeaves the question out of the session's total points. It's still graded.
    templateIdstringNoThe workspace rubric template this rubric was created from, if any.

    Each criterion, and each level of an analytical criterion, has these fields:

    FieldTypeRequiredDescription
    idstringYesA unique ID. You'll find it again in the grading results, so pick something you can recognise.
    titlestringanalytical, holisticA short name. Used for analytical criteria and levels, and holistic criteria.
    descriptionstringsimple, holistic, exact-valuesWhat an answer needs to earn the points. For exact-values, the accepted value itself.
    pointsnumberYesThe points awarded when the criterion or level applies.
    minPointsnumberNoOptional lower bound, when the criterion or level allows a range of points.
    maxPointsnumberNoOptional upper bound for that range.
    levelsarrayNoOnly for analytical criteria: the levels to choose from, each with the fields above.
    Rubrics are normalised when saved

    Examplary tidies up rubrics when you save a question: missing ids are generated, fields that don't apply to the rubric type are dropped (such as levels outside analytical rubrics, or title on simple criteria), and analytical levels and holistic criteria are sorted by points. Read the question back after saving if you need the final shape.

    Simple

    A checklist: each criterion describes something the answer should contain, and every criterion that applies adds its points. Use it for answers that need to cover several separate points.

    scoring (simple)
    {
      "rubricType": "simple",
      "criteria": [
        {
          "id": "mentions-42",
          "description": "Mentions the number 42.",
          "points": 1
        },
        {
          "id": "names-book",
          "description": "Names The Hitchhiker's Guide to the Galaxy.",
          "points": 1
        },
        {
          "id": "explains-joke",
          "description": "Explains why the answer is a joke.",
          "points": 2
        }
      ]
    }

    The maximum is the sum of all criteria, 4 points here. In the results, every criterion gets its own entry with the points the AI awarded for it:

    Answer (simple)
    {
      "gradingStatus": "auto-graded",
      "gradedBy": "system-ai",
      "pointsAwarded": 2,
      "feedback": "<p>You found the answer, but didn't say where it comes from.</p>",
      "scoringCriteria": [
        {
          "id": "mentions-42",
          "pointsAwarded": 1,
          "aiPointsAwarded": 1,
          "confidenceScore": 95,
          "reasoning": "States 42 directly."
        },
        {
          "id": "names-book",
          "pointsAwarded": 0,
          "aiPointsAwarded": 0,
          "confidenceScore": 90,
          "reasoning": "No reference to the book."
        },
        {
          "id": "explains-joke",
          "pointsAwarded": 1,
          "aiPointsAwarded": 1,
          "confidenceScore": 60,
          "reasoning": "Hints at the joke but doesn't explain it."
        }
      ]
    }

    Analytical

    A grid: each criterion is a dimension of the answer, such as correctness or argumentation, with levels from weak to strong. The AI picks exactly one level per criterion, and the levels' points add up to the score. Analytical criteria only have a title and levels; the points live on the levels.

    scoring (analytical)
    {
      "rubricType": "analytical",
      "criteria": [
        {
          "id": "correctness",
          "title": "Correctness",
          "levels": [
            {
              "id": "correctness-0",
              "title": "Incorrect",
              "description": "Doesn't mention 42.",
              "points": 0
            },
            {
              "id": "correctness-1",
              "title": "Partially correct",
              "description": "Mentions 42, but not the book.",
              "points": 1
            },
            {
              "id": "correctness-2",
              "title": "Correct",
              "description": "Mentions 42 and the book.",
              "points": 2
            }
          ]
        },
        {
          "id": "clarity",
          "title": "Clarity",
          "levels": [
            { "id": "clarity-0", "title": "Unclear", "points": 0 },
            { "id": "clarity-1", "title": "Clear", "points": 1 }
          ]
        }
      ]
    }

    The maximum is the sum of each criterion's highest level, 3 points here. In the results, each criterion's entry also says which level was picked, in selectedLevelId:

    Answer (analytical)
    {
      "gradingStatus": "auto-graded",
      "gradedBy": "system-ai",
      "pointsAwarded": 2,
      "feedback": "<p>Correct number! Mention where it comes from for full marks.</p>",
      "scoringCriteria": [
        {
          "id": "correctness",
          "selectedLevelId": "correctness-1",
          "aiSelectedLevel": "correctness-1",
          "pointsAwarded": 1,
          "aiPointsAwarded": 1,
          "confidenceScore": 85,
          "reasoning": "Mentions 42, but not the book."
        },
        {
          "id": "clarity",
          "selectedLevelId": "clarity-1",
          "aiSelectedLevel": "clarity-1",
          "pointsAwarded": 1,
          "aiPointsAwarded": 1,
          "confidenceScore": 90,
          "reasoning": "A short, direct answer."
        }
      ]
    }

    Holistic

    A single scale for the answer as a whole: each criterion describes one overall quality, and the AI picks the one that fits best. Use it when an answer is best judged in one go, like a short essay.

    scoring (holistic)
    {
      "rubricType": "holistic",
      "criteria": [
        {
          "id": "insufficient",
          "title": "Insufficient",
          "description": "Misses the point of the question.",
          "points": 0
        },
        {
          "id": "adequate",
          "title": "Adequate",
          "description": "Gives the right answer without context.",
          "points": 2
        },
        {
          "id": "excellent",
          "title": "Excellent",
          "description": "Gives the answer and explains its origin.",
          "points": 4
        }
      ]
    }

    The maximum is the highest criterion, 4 points here. In the results, every criterion has an entry, but only the picked one has points; the others are 0:

    Answer (holistic)
    {
      "gradingStatus": "auto-graded",
      "gradedBy": "system-ai",
      "pointsAwarded": 2,
      "scoringCriteria": [
        {
          "id": "insufficient",
          "pointsAwarded": 0,
          "aiPointsAwarded": 0,
          "confidenceScore": 90,
          "reasoning": "…"
        },
        {
          "id": "adequate",
          "pointsAwarded": 2,
          "aiPointsAwarded": 2,
          "confidenceScore": 80,
          "reasoning": "Right answer, no context."
        },
        {
          "id": "excellent",
          "pointsAwarded": 0,
          "aiPointsAwarded": 0,
          "confidenceScore": 85,
          "reasoning": "…"
        }
      ]
    }

    To find the picked criterion, look for the entry with points. If a teacher changes the grade in Examplary, the results instead hold just the picked criterion, with selectedLevelId set to its own id.

    Exact values

    For answers with a known correct value, like fill-in-the-blank questions. Each criterion's description is an accepted value, with the points it's worth. There's no AI involved: the answer is graded straight away when it's saved.

    scoring (exact-values)
    {
      "rubricType": "exact-values",
      "exactValuesCaseInsensitive": true,
      "criteria": [
        { "id": "full", "description": "42", "points": 2 },
        { "id": "written-out", "description": "forty-two", "points": 1 }
      ]
    }

    The first matching criterion wins, and the maximum is the highest criterion, 2 points here. Answers are compared after trimming surrounding whitespace. With exactValuesCaseInsensitive, capitalisation and all whitespace are ignored too, so "New York" matches "newyork". See exact-value matching for the details.

    The results are final straight away, and hold the matching criterion only:

    Answer (exact-values)
    {
      "gradingStatus": "graded",
      "gradedBy": "system-exact-values",
      "pointsAwarded": 2,
      "scoringCriteria": [
        { "id": "full", "selectedLevelId": "full", "pointsAwarded": 2 }
      ]
    }

    An answer that matches nothing gets "pointsAwarded": 0 and an empty scoringCriteria.

    Reading the results

    Grading results are part of each answer in the assessment session, which you get from the sessions endpoints or the exam.session.grading.completed webhook. See step 4 of the developer guide for how to fetch them.

    FieldDescription
    gradingStatusauto-grading-pending, auto-graded (an AI suggestion, not yet reviewed), auto-grading-failed, graded (final), or not-graded.
    gradedBysystem-ai, system-exact-values, system-response-processing (graded by the question type), or the ID of the teacher who graded it.
    pointsAwardedThe answer's points, rounded to one decimal.
    pointsToEarnThe question's maximum points.
    feedbackFeedback for the student, as HTML (see below). Not always present.
    scoringCriteriaThe score per criterion, matched to your rubric by id. See below.
    gradedAtWhen the answer was graded.

    Each entry in scoringCriteria has:

    FieldDescription
    idThe criterion's id from the rubric.
    pointsAwardedThe points for this criterion.
    selectedLevelIdThe picked level, for analytical criteria (and the matched value, for exact values).
    reasoningThe AI's explanation of the score, as HTML (see below). Useful to show to teachers next to the grade.
    confidenceScoreHow sure the AI is, from 0 to 100.
    aiPointsAwardedWhat the AI originally awarded. Stays the same when a teacher changes pointsAwarded.
    aiSelectedLevelThe level the AI originally picked, for analytical criteria.

    Both feedback and reasoning can contain HTML formatting, plus two custom tags: <inline-math> for formulas, and <content-reference> to quote part of the student's answer. Its quote attribute holds the exact text it refers to:

    The student mentioned <content-reference quote="42">"42"</content-reference> but
    did not name the book.

    When showing these in your own interface, you can highlight the quoted text in the answer, or just render the tag's content as plain text.

    The AI's grades are suggestions, so they come back as auto-graded and count towards the session's points but not its final grade. A teacher can review them in Examplary, or you can finalise them yourself: accept a suggestion as it is, accept all of a session's suggestions, or save your own grade. Either way the answer becomes graded, and a late AI result never overwrites it.

    Getting better grades from the AI

    Clear, specific criterion and level descriptions make the biggest difference. Add guidance for anything that applies across criteria, and a modelAnswer to show what full marks looks like. To teach the AI your own grading style, add previously graded answers as examples.