Writing rubrics for Reasonate
The rubric is the most important thing you'll write. It's a checklist of attributes, the specific things a student must demonstrate understanding of, and it directly determines their Mastery Level.
The core principle
A Reasonate rubric is not a list of acceptable answers. It's a list of what a student must demonstrate understanding of, the concepts, mechanisms, distinctions, and reasoning chains they need to articulate to prove they actually get it.
Write your rubric the way you'd brief a TA who's going to verbally examine a student: “make sure they explain X, get them to articulate Y, watch for the common mistake of Z.”
How a rubric becomes a Mastery Level
Reasonate doesn't ask an AI to “grade” the chat. After the conversation, an extractor reads the whole transcript and, for each attribute in your rubric, records one of three things: the student demonstrated it on their own, demonstrated it only with a hint (scaffolded), or didn't show it. A fixed rule then turns those per-attribute results into a Mastery Level, the same rule for every student, every time.
Here is the ladder your rubric drives. The jump to L4–L5 is grounded in the Feynman technique: it's about defending understanding under pushback and explaining it clearly, not about applying it to a new case.
Novel transfer (applying the idea to a brand-new case) is no longer required for any level, it still counts when it happens naturally, but the AI won't force a cold “now apply this somewhere else” pivot.
Evidence Confidence (paste / tab-away / typing checks) is a separate measure, it can withhold XP, but it never lowers the Mastery Level. The level reflects only what was demonstrated.
The six attribute types
Every rubric is built from six attribute categories, each with a fixed role in the ladder above. Aim to write at least one strong attribute for each, generated rubrics already include all six, so you're usually refining what each one probes, not inventing them.
The core concept(s) stated correctly and completely. Write the key terms, claims, or factors the student must name and define in their own words.
A concrete, valid example the student walks through themselves. Write the kind of case or instance that should count, and what makes an example actually demonstrate the idea.
The why/how, the causal chain or process, not just the what. Write the mechanism steps the student must connect (“X happens because Y, which drives Z”).
Applying the idea to a NEW case. Still valuable when it happens naturally, and you can keep it in a rubric, but it is no longer required for L4 or any level, and the AI will not force a cold transfer pivot.
The most important thing to write well. The AI floats your critical misconceptions back at the student as plausible-but-wrong ideas; holding firm and explaining why they're wrong is how a student earns L4. Write them under Critical misconceptions and make them GENUINELY plausible (a real wrong intuition, not a strawman). An unresolved one still caps the score at L2.
Knowing the limits, where the idea stops applying, edge cases, or a competing view. A contributing L5 signal. Write the boundary or counterclaim the student should be able to articulate.
Defense & clarity, what L4 and L5 reward
Advanced mastery is the Feynman test: can the student defend the idea under pushback, and explain it simply?
Defense (L4). After the student shows the core, the AI, staying in its curious-student character, floats one of your critical misconceptions as a plausible-but-wrong idea, built on the student's own words (“Wait, you said X, wouldn't that mean [misconception]?”). Holding firm and explaining why it's wrong earns the defense; caving does not. This is why your critical misconceptions are now the highest-leverage thing you write: a vague or strawman misconception makes a weak challenge, while a genuinely plausible one makes a real test of understanding. The AI only ever floats misconceptions you listed in the rubric, so write them well.
Clarity (L5). L5 also requires the student to explain the topic clearly, plain language a motivated non-expert could follow, jargon defined when used. It is judged holistically over the whole conversation; there is no awkward “explain it to a child” step. Writing attributes that push for plain, mechanism-level explanation (not jargon recitation) is what surfaces clarity.
Discrimination is conditional: if a student says something actually wrong and the AI corrects them, revising gracefully is good and stubbornly resisting a correct correction is the failure mode. But a flawless student who never errs still reaches L5, they are never required to produce a “revise” moment they had no occasion for.
The three descriptors per attribute
For each attribute you write three short descriptions of what a student's answer looks like at each level of evidence. These guide how the extractor labels the transcript, so be concrete:
Why this matters: independent vs. with-help determines how solidly each foundational attribute counts toward L2–L3 and how polished the score looks. (L4 and L5 are decided separately, by defense and clarity, see above.) The clearer your descriptors, the more consistent the labeling.
What makes a good rubric
- Specific concepts, mechanisms, or causal chains the student must articulate
- GENUINELY PLAUSIBLE misconceptions, the AI floats these as wrong-but-tempting challenges, so they directly drive L4 defense
- Distinctions the student must make (X vs Y) where confusion is common
- Attributes that push for plain, mechanism-level explanation (this is what surfaces L5 clarity)
- Vague terms like 'good understanding' or 'demonstrates mastery'
- Lists of acceptable answer phrasings (the AI doesn't grade phrasing)
- Rules about length, format, or grammar
- Pure factual recall ('list the 5 steps'), Reasonate grades reasoning, not memory
Strong example vs. weak example
“Student should understand photosynthesis well and explain the main concepts clearly.”
Why: Doesn't tell the AI what specifically to probe. Every learner will get probed differently, which makes scores inconsistent.
“The learner must explain: (1) the inputs (CO2, water, light) and outputs (glucose, O2); (2) the two stages, light-dependent reactions in the thylakoid, and the Calvin cycle in the stroma; (3) the role of ATP and NADPH as energy carriers between stages, in plain language. Probe whether they understand mechanism, not just memorize inputs/outputs. Critical misconception to challenge them with: 'plants get their food/mass from the soil', a genuinely common wrong intuition they should hold firm against.”
Why: Maps cleanly onto the attributes (definition, example, mechanism), pushes for a plain-language explanation (clarity), and supplies a plausible misconception the AI can float as a defense challenge, the path to L4/L5.
Templates you can copy
The learner must explain: (1) the inputs and outputs; (2) the key stages and what happens in each; (3) the mechanism by which inputs become outputs; (4) WHERE this happens (cellular, physical, conceptual location); (5) at least one concrete example. Push for mechanism, surface answers (“X makes Y”) should get followed up with “but HOW?”
The learner must explain: (1) the main causal factors; (2) how those factors link together to produce the outcome (the chain, not just the list); (3) why this outcome happened HERE and THEN, what made the conditions ripe; (4) at least one counterfactual (“what would have prevented this?”); (5) the most common misconception about this topic.
The learner must explain: (1) the formal definition of each concept; (2) where the two concepts overlap (and why people confuse them); (3) the precise dividing line; (4) a concrete example of each that the other does NOT cover; (5) why the distinction matters in practice.
The learner must explain: (1) what each term in the formula represents physically/conceptually; (2) the derivation or intuition (not just the formula); (3) when the formula applies vs. when it breaks; (4) a worked example; (5) the most common error students make using this formula.
A rubric-writing checklist
- Does the rubric cover definition, example, and mechanism (the L2–L3 backbone)?
- Does it list GENUINELY PLAUSIBLE critical misconceptions for the AI to challenge with? (These drive L4 defense, the highest-leverage thing to write.)
- Does it push for plain, mechanism-level explanation, not jargon recitation? (This surfaces L5 clarity.)
- Does it focus on reasoning, not just facts?
- Could the AI use this rubric to challenge and probe without you in the room?
- Would two different graders following this rubric arrive at similar levels?
