How to write an evaluation rubric

This guide to evaluation rubric writing shows you how to turn "looks good to me" into criteria a grader can actually apply. A rubric is a contract with your grader: it says exactly what counts as right, what counts as wrong, and how wrong is how bad.

We grade outputs. We do not train models, run fine-tuning, or sell experts we cannot verify.

The short version

What is an evaluation rubric?

An evaluation rubric is the written standard a grader uses to judge model outputs. It names the criteria, describes what each level of quality looks like, and gives examples so two graders reach the same verdict on the same sample. Without one, every grader invents their own scale and your results are noise. With one, grading becomes repeatable, and repeatable grading is the whole point of evaluation.

Where it breaks

What bad rubrics do to your results

A bad rubric does not just slow grading down. It manufactures disagreement.

Failure pattern

The 1-to-5 shrug

"Rate quality from 1 to 5." Every grader has a different 3. One person's 4 is another's 2, and the average of their guesses tells you nothing. A number without a description is not a measurement.

Failure pattern

Criteria that overlap

"Accuracy," "correctness," and "factual quality" as three separate criteria. Graders split the same judgment three ways or triple-count one flaw. Overlapping criteria inflate some failures and hide others.

Failure pattern

Levels nobody can tell apart

Five levels where levels 2, 3, and 4 all read like "kind of okay." If a grader cannot reliably pick between adjacent levels, the scale is decoration. Fewer, sharper levels beat more, blurrier ones.

Failure pattern

No examples

Criteria described in the abstract, with zero sample outputs showing what each level looks like in practice. Graders then calibrate on the first few samples they see, which means the rubric drifts as grading goes on.

The method

Anatomy of a good rubric

A good rubric has four parts. Miss any one of them and graders start improvising.

1. Named criteria, each testing one thing

List the dimensions you grade: factual accuracy, instruction following, tone, safety, whatever matters for your product. Each criterion should test exactly one thing, with no overlap. Three to five criteria is the sweet spot for most outputs. More than that and graders slow down and start skimming.

2. Levels with plain descriptions

For each criterion, describe what each level looks like in words a new grader understands on first read. We use four severity levels everywhere: P0 Critical, P1 Major, P2 Minor, P3 Clean. Our severity scale page shows the exact wording. Four levels is enough to separate "fails" from "needs work" from "fine," and few enough that graders pick consistently.

3. Examples for every level

One real or realistic sample output per level, per criterion, annotated with why it earned that grade. Examples do more calibration work than paragraphs of description. When two graders disagree, the examples are the referee: "this sample looks like the P1 example, not the P2 one."

4. Edge-case rules

The calls you know will come up: what if the answer is right but in the wrong language? What if it refuses? What if it is correct but twice the length limit? Write the ruling down before grading starts. Every edge case decided during grading is a small inconsistency injected into your results.

Worked example

One rubric row, fully written

Say you are grading a support chatbot's answers. Here is one criterion, "Factual accuracy," written out the way we would write it. Notice each level has a description and a concrete example.

LevelWhat it meansExample
P0 CriticalThe answer states something false about the product with confidence."Your plan includes unlimited international calls." The plan does not. A customer acting on this gets a surprise bill.
P1 MajorThe answer is materially misleading or incomplete in a way that changes what the customer should do."You can cancel anytime from settings." Cancellation actually requires contacting support. The customer will try, fail, and complain.
P2 MinorSmall inaccuracy that does not change the meaning or the customer's next step."Our support team replies within a couple of hours." The documented target is one business day. Slightly off, same action.
P3 CleanEvery checkable claim matches the source material."You can upgrade from the billing page under Settings." Matches the help article exactly.

How a grader uses this row

The grader reads the sample, checks each claim against the help articles, and matches what they see to a row in the table. "Unlimited international calls" matches the P0 description, so the sample gets P0 on this criterion with the example quoted as the reason. No judgment call, no improvisation. That is what a rubric is for.

Adapting it to your product

Swap the criterion for yours, keep the structure: level, plain description, concrete example. Write three to five rows like this one and you have a rubric. If writing it feels hard, that is useful information: it means the standard was fuzzy in your head, and fuzzy standards produce fuzzy products. Our pilot includes rubric help for exactly this reason.

How it works

From samples to answers in three steps

01

Send samples

You send 20 to 50 model outputs plus your rubric, or we help you write one.

02

We grade

Trained reviewers grade every sample against the rubric, backed by automated checks.

03

You get the report

Sample-level grades, issue patterns, recommended fixes, and a walkthrough call.

Deliverables

What you get

  • Sample-level gradesEvery sample scored P0 to P3 with a written reason.
  • Issue summaryThe patterns across the batch, ranked by severity and frequency.
  • Recommended fixesConcrete next steps for your prompts, retrieval, or guardrails.
  • Walkthrough callWe go through the report with your team and answer questions.
The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

Three to five for most outputs. Fewer and you miss real dimensions; more and graders slow down and start skimming, which hurts consistency. Start with the failures that would actually hurt you and write one criterion per failure type.

Usually not. More levels feel precise but graders cannot reliably tell them apart, so the extra levels just add noise. Four levels, fails, needs work, minor issue, clean, cover the decisions you actually make about an output. That is why we grade everything P0 to P3.

Someone who knows what good looks like for your product, ideally the person who would be upset by a bad output reaching users. We can help: every pilot includes rubric review, and we will pressure-test your criteria against real samples before grading starts.

Sometimes, if the outputs are close cousins like emails and chat replies. If the outputs differ a lot, say code and marketing copy, write separate rubrics. A rubric stretched across unlike outputs gets vague, and vague rubrics grade nothing well.

Have two people grade the same ten samples independently, then compare. If they agree on most grades, the rubric is working. If they argue, the argument tells you exactly which criterion or level needs rewriting. Do this before the full grading run, not after.

Yes. Send it as-is and we will grade against it, or send a draft and we will tighten it with you first. Either way, you approve the final rubric before any grading begins. See our sample report for what graded output looks like.

A rubric is a contract. Write a good one.

Send us your draft or your samples. We will help you tighten it, then grade against it.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.