This guide to evaluation rubric writing shows you how to turn "looks good to me" into criteria a grader can actually apply. A rubric is a contract with your grader: it says exactly what counts as right, what counts as wrong, and how wrong is how bad.
We grade outputs. We do not train models, run fine-tuning, or sell experts we cannot verify.
An evaluation rubric is the written standard a grader uses to judge model outputs. It names the criteria, describes what each level of quality looks like, and gives examples so two graders reach the same verdict on the same sample. Without one, every grader invents their own scale and your results are noise. With one, grading becomes repeatable, and repeatable grading is the whole point of evaluation.
A bad rubric does not just slow grading down. It manufactures disagreement.
"Rate quality from 1 to 5." Every grader has a different 3. One person's 4 is another's 2, and the average of their guesses tells you nothing. A number without a description is not a measurement.
"Accuracy," "correctness," and "factual quality" as three separate criteria. Graders split the same judgment three ways or triple-count one flaw. Overlapping criteria inflate some failures and hide others.
Five levels where levels 2, 3, and 4 all read like "kind of okay." If a grader cannot reliably pick between adjacent levels, the scale is decoration. Fewer, sharper levels beat more, blurrier ones.
Criteria described in the abstract, with zero sample outputs showing what each level looks like in practice. Graders then calibrate on the first few samples they see, which means the rubric drifts as grading goes on.
A good rubric has four parts. Miss any one of them and graders start improvising.
List the dimensions you grade: factual accuracy, instruction following, tone, safety, whatever matters for your product. Each criterion should test exactly one thing, with no overlap. Three to five criteria is the sweet spot for most outputs. More than that and graders slow down and start skimming.
For each criterion, describe what each level looks like in words a new grader understands on first read. We use four severity levels everywhere: P0 Critical, P1 Major, P2 Minor, P3 Clean. Our severity scale page shows the exact wording. Four levels is enough to separate "fails" from "needs work" from "fine," and few enough that graders pick consistently.
One real or realistic sample output per level, per criterion, annotated with why it earned that grade. Examples do more calibration work than paragraphs of description. When two graders disagree, the examples are the referee: "this sample looks like the P1 example, not the P2 one."
The calls you know will come up: what if the answer is right but in the wrong language? What if it refuses? What if it is correct but twice the length limit? Write the ruling down before grading starts. Every edge case decided during grading is a small inconsistency injected into your results.
Say you are grading a support chatbot's answers. Here is one criterion, "Factual accuracy," written out the way we would write it. Notice each level has a description and a concrete example.
| Level | What it means | Example |
|---|---|---|
| P0 Critical | The answer states something false about the product with confidence. | "Your plan includes unlimited international calls." The plan does not. A customer acting on this gets a surprise bill. |
| P1 Major | The answer is materially misleading or incomplete in a way that changes what the customer should do. | "You can cancel anytime from settings." Cancellation actually requires contacting support. The customer will try, fail, and complain. |
| P2 Minor | Small inaccuracy that does not change the meaning or the customer's next step. | "Our support team replies within a couple of hours." The documented target is one business day. Slightly off, same action. |
| P3 Clean | Every checkable claim matches the source material. | "You can upgrade from the billing page under Settings." Matches the help article exactly. |
The grader reads the sample, checks each claim against the help articles, and matches what they see to a row in the table. "Unlimited international calls" matches the P0 description, so the sample gets P0 on this criterion with the example quoted as the reason. No judgment call, no improvisation. That is what a rubric is for.
Swap the criterion for yours, keep the structure: level, plain description, concrete example. Write three to five rows like this one and you have a rubric. If writing it feels hard, that is useful information: it means the standard was fuzzy in your head, and fuzzy standards produce fuzzy products. Our pilot includes rubric help for exactly this reason.
You send 20 to 50 model outputs plus your rubric, or we help you write one.
Trained reviewers grade every sample against the rubric, backed by automated checks.
Sample-level grades, issue patterns, recommended fixes, and a walkthrough call.
Three to five for most outputs. Fewer and you miss real dimensions; more and graders slow down and start skimming, which hurts consistency. Start with the failures that would actually hurt you and write one criterion per failure type.
Usually not. More levels feel precise but graders cannot reliably tell them apart, so the extra levels just add noise. Four levels, fails, needs work, minor issue, clean, cover the decisions you actually make about an output. That is why we grade everything P0 to P3.
Someone who knows what good looks like for your product, ideally the person who would be upset by a bad output reaching users. We can help: every pilot includes rubric review, and we will pressure-test your criteria against real samples before grading starts.
Sometimes, if the outputs are close cousins like emails and chat replies. If the outputs differ a lot, say code and marketing copy, write separate rubrics. A rubric stretched across unlike outputs gets vague, and vague rubrics grade nothing well.
Have two people grade the same ten samples independently, then compare. If they agree on most grades, the rubric is working. If they argue, the argument tells you exactly which criterion or level needs rewriting. Do this before the full grading run, not after.
Yes. Send it as-is and we will grade against it, or send a draft and we will tighten it with you first. Either way, you approve the final rubric before any grading begins. See our sample report for what graded output looks like.
Send us your draft or your samples. We will help you tighten it, then grade against it.
Not sure where to start? Talk to us and we will point you at the right evaluation.
Lock a quick 15-minute intro call — we'll scope your evaluation needs and deploy vetted experts within 48 hours.