Legal AI cannot afford confident errors.

A fabricated case citation in a brief is not a typo, it is a sanction risk. Legal AI evaluation grades your system's outputs against your sources and standards, citation by citation.

We grade outputs. We grade your AI's outputs against your rubric and source materials. We do not provide lawyers or licensed professionals, and our grading does not replace professional legal review.

The short version

What is legal AI evaluation?

Legal AI output evaluation is the strict grading of AI-generated legal work, drafts, summaries, research memos, contract markups, against your sources and your standards. Legal tech teams need it because the failure mode is uniquely dangerous: a fluent memo with a fabricated citation looks exactly like good work until a judge checks it. Good looks like every citation verified against a real source, no invented holdings, and language that stays inside the system's approved role. Plainly: we grade outputs. We are not lawyers and our reports are not legal advice.

Where it breaks

The failures that end up in front of a judge

Legal AI fails in ways that look like competence. We grade for the patterns that matter.

Fabricated citations

The case that never existed

The most reported legal AI failure: a real-looking citation to a case that does not exist, or a real case cited for a holding it never made. We verify citations against your source materials and flag every one that does not hold up.

Overstated authority

Dicta dressed up as holdings

Subtler than invention: the AI cites a real case but upgrades what it said. A passing remark becomes the rule of the case. We grade whether the output's claims match what the source actually supports, not just whether the source is real.

Jurisdiction drift

Right law, wrong place

Legal answers are jurisdiction-specific, and models mix them freely. A memo about California procedure citing New York rules reads fine to anyone who is not checking. We grade jurisdiction consistency against the matter's actual venue.

Advice beyond scope

Drafting that sounds like counsel

Many legal AI tools are meant to assist, not advise. When the output crosses into direct legal advice without the right framing, that is a product risk. We grade whether outputs stayed inside their approved role.

How it works

From samples to answers in three steps

01

Send samples

You send 20 to 50 outputs plus your source materials and standards, or we help you write a rubric from them. Citations get verified against the sources you provide.

02

We grade

Trained reviewers grade every sample against your rubric, backed by automated checks. Fabricated citations and overstated authority are treated as the P0s they are.

03

You get the report

Sample-level grades, citation verification results, the failure patterns across the batch, recommended fixes, and a walkthrough call.

Deliverables

What you get

  • Sample-level gradesEvery output scored P0 to P3 with a written reason.
  • Citation verificationEach citation checked against your sources: real, misused, or invented.
  • Authority accuracy checkWhether claims match what the cited sources actually support.
  • Recommended fixesConcrete next steps for your prompts, retrieval, and review workflow.
How we grade

How we grade legal outputs

Citations are verified, not trusted

Every citation in every sample is checked against the sources you provide. Does it exist. Does it say what the output claims. Does it support the point it is attached to. A real case cited for a holding it never made fails just as hard as an invented one, because in front of a judge the difference does not matter. Our human evaluation process is built for exactly this kind of careful, line-by-line work.

Authority gets graded in degrees

Legal writing constantly calibrates how strongly it states things: holdings versus dicta, majority versus concurrence, binding versus persuasive. We grade whether the output's confidence matches its sources. An output that upgrades a passing remark into the rule of the case is overstating authority, and the report flags it as its own pattern, distinct from outright invention.

We also grade consistency across the batch. If the same clause gets summarized three different ways in three samples, that variance is a pattern worth knowing before opposing counsel finds it. The report flags inconsistent handling of repeated task types so you can tighten the prompt or the template. Consistency matters in legal work because downstream readers assume the system treats like cases alike. When it does not, every output becomes suspect, even the correct ones. A system that is right in five different ways is harder to trust than one that is right the same way every time.

The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

No. We grade AI outputs against the rubric and sources you provide. We are not lawyers, we do not employ licensed attorneys as graders, and our reports are not legal advice and do not replace review by qualified counsel. What you get is a documented, systematic read on where your system's outputs break your own standards.

Against the source materials you provide: your case database, your document set, your research corpus. We check that each citation exists, that it says what the output claims it says, and that it supports the point it is attached to. A real case cited for the wrong holding still fails.

Research memos, brief drafts, contract markup suggestions, clause summaries, deposition summaries, and similar text outputs. If your team can define what a correct output looks like and provide the sources, we can grade against them.

Yes, when you tell us the jurisdiction. The rubric includes a jurisdiction check: the output's rules, procedures, and citations must match the matter's actual venue. Mixed-jurisdiction answers are one of the most common failures we see in legal AI.

Attorney review is essential and expensive. We give you the systematic layer underneath it: every sample graded the same way, citation problems counted, failure patterns named. That lets your attorneys spend their time on judgment calls instead of catching the same fabricated citation repeatedly.

Hallucination detection asks whether outputs contain invented facts. Legal AI evaluation goes further: it checks whether real sources are used correctly, whether authority is overstated, and whether the output stayed in its approved role. For legal work, correct use of real sources matters as much as catching inventions.

Check every citation before a judge does.

Send 20 to 50 outputs with your sources. We will grade them all.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.