How to Read an AI Grading Report: Your Report, Decoded

An ai grading report is only useful if you can read it. This guide walks through every section of our report format: what the severities mean, how to read a sample-level grade, and how to turn issue patterns into fixes.

We grade outputs. We do not train models, run fine-tuning, or sell experts we cannot verify.

The short version

What is an AI grading report?

An AI grading report is the document you get after an evaluation: every sampled output graded against your rubric, the failure patterns across the batch, and recommended fixes ranked by impact. It is the bridge between "the model feels off" and "here is exactly what to change."

A good report reads like an engineering document, not a dashboard. Each grade comes with a written reason and a suggested fix, so your team can act without scheduling a meeting to interpret the numbers. If you want to see the shape of one before committing, look at our sample report.

Reading mistakes

Four ways teams misread grading reports

A report is only as good as its reading. These are the misreadings we see most often, and how to avoid each one.

Averages

Averages hide the P0s

A batch can average a respectable score while containing five critical failures. Always read the severity distribution first: how many P0s, how many P1s. The average tells you the weather; the P0 count tells you whether the roof is on fire.

Skimming

Reading the summary, skipping the samples

The issue summary is a map, not the territory. Pick three graded samples and read them end to end: the output, the grade, the reason, the fix. That is where the report stops being abstract and starts being obvious.

Treating patterns as one-offs

One bug, fifty samples

When the same failure shows up across many samples, it is one bug with fifty symptoms, not fifty separate problems. Fix the pattern at its source, usually the prompt, the retrieval, or a guardrail, and the whole cluster clears at once.

Filing it away

The report nobody acted on

A grading report has a half-life. Read it within days, assign the fixes, and schedule the re-grade. Reports that sit for a month describe a system that no longer exists. The walkthrough call exists precisely to turn reading into a fix list.

How it works

From samples to answers in three steps

01

Send samples

You send 20 to 50 model outputs plus your rubric, or we help you write one.

02

We grade

Trained reviewers grade every sample against the rubric, backed by automated checks.

03

You get the report

Sample-level grades, issue patterns, recommended fixes, and a walkthrough call.

Deliverables

What you get

  • Sample-level gradesEvery sample scored P0 to P3 with a written reason.
  • Issue summaryThe patterns across the batch, ranked by severity and frequency.
  • Recommended fixesConcrete next steps for your prompts, retrieval, or guardrails.
  • Walkthrough callWe go through the report with your team and answer questions.
Reading the report

A walkthrough of each section

Start with the P0s

Open the report at the severity distribution. Every P0 is a sample that failed outright: a fabricated fact, unsafe content, or a completely wrong answer. Read each P0 row fully before anything else, because these are the outputs that would have hurt you in production.

For each P0, note the reason and the suggested fix. P0s usually cluster into one or two patterns, which means a small number of fixes clears most of them. That clustering is the most valuable page in the report.

Read three sample rows end to end

Pick three samples across severities and read the full row: the output, the grade, the written reason, the suggested fix. This teaches you how the rubric was applied better than any methodology section. You will also spot whether any grade feels wrong, which is worth raising on the walkthrough call.

Pay attention to the P1s. They are the judgment calls: materially wrong but not catastrophic. How your team reacts to the P1s tells you where your quality bar actually sits. Our severity scale explains the reasoning behind each level.

Bring the patterns to your prompt review

The issue summary ranks patterns by severity and frequency. Take the top three patterns into your next prompt or pipeline review. Each pattern maps to a concrete change: a prompt instruction, a retrieval fix, a guardrail, or a fallback.

After the fixes ship, re-grade. The second report is where evaluation pays for itself: you see the P0 count drop, or you learn the fix did not work and try the next one. Improvement you can measure beats improvement you can feel.

The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

P0 means the sample failed outright: fabricated facts presented confidently, unsafe content, or a completely wrong answer. It is the severity that would cause real harm or embarrassment in production. Every P0 in your report deserves a fix before the related feature ships.

Because grading is about correctness, not fluency. An answer can read beautifully and still be materially wrong or misleading. That gap, between sounding right and being right, is exactly what grading catches and what casual review misses.

Take the top-ranked patterns into your prompt or pipeline review. Each pattern points at a source: the prompt wording, the retrieved context, a missing guardrail. Fix the source, not the samples. Then re-grade to confirm the pattern actually cleared.

Yes. Reports are written to be shared: plain language, no jargon walls, every grade explained. Product managers read the summary, engineers read the sample rows and fixes. If part of it is unclear to a non-specialist, tell us on the walkthrough call and we will clarify it.

We keep what we need to deliver the work and answer follow-up questions, and we handle your samples as your confidential material. If you have specific data handling requirements, tell us before the pilot starts and we will work within them.

Get a report your team can actually read

Send samples and see what a graded batch looks like.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.