A fabricated case citation in a brief is not a typo, it is a sanction risk. Legal AI evaluation grades your system's outputs against your sources and standards, citation by citation.
We grade outputs. We grade your AI's outputs against your rubric and source materials. We do not provide lawyers or licensed professionals, and our grading does not replace professional legal review.
Legal AI output evaluation is the strict grading of AI-generated legal work, drafts, summaries, research memos, contract markups, against your sources and your standards. Legal tech teams need it because the failure mode is uniquely dangerous: a fluent memo with a fabricated citation looks exactly like good work until a judge checks it. Good looks like every citation verified against a real source, no invented holdings, and language that stays inside the system's approved role. Plainly: we grade outputs. We are not lawyers and our reports are not legal advice.
Legal AI fails in ways that look like competence. We grade for the patterns that matter.
The most reported legal AI failure: a real-looking citation to a case that does not exist, or a real case cited for a holding it never made. We verify citations against your source materials and flag every one that does not hold up.
Subtler than invention: the AI cites a real case but upgrades what it said. A passing remark becomes the rule of the case. We grade whether the output's claims match what the source actually supports, not just whether the source is real.
Legal answers are jurisdiction-specific, and models mix them freely. A memo about California procedure citing New York rules reads fine to anyone who is not checking. We grade jurisdiction consistency against the matter's actual venue.
Many legal AI tools are meant to assist, not advise. When the output crosses into direct legal advice without the right framing, that is a product risk. We grade whether outputs stayed inside their approved role.
You send 20 to 50 outputs plus your source materials and standards, or we help you write a rubric from them. Citations get verified against the sources you provide.
Trained reviewers grade every sample against your rubric, backed by automated checks. Fabricated citations and overstated authority are treated as the P0s they are.
Sample-level grades, citation verification results, the failure patterns across the batch, recommended fixes, and a walkthrough call.
Every citation in every sample is checked against the sources you provide. Does it exist. Does it say what the output claims. Does it support the point it is attached to. A real case cited for a holding it never made fails just as hard as an invented one, because in front of a judge the difference does not matter. Our human evaluation process is built for exactly this kind of careful, line-by-line work.
Legal writing constantly calibrates how strongly it states things: holdings versus dicta, majority versus concurrence, binding versus persuasive. We grade whether the output's confidence matches its sources. An output that upgrades a passing remark into the rule of the case is overstating authority, and the report flags it as its own pattern, distinct from outright invention.
We also grade consistency across the batch. If the same clause gets summarized three different ways in three samples, that variance is a pattern worth knowing before opposing counsel finds it. The report flags inconsistent handling of repeated task types so you can tighten the prompt or the template. Consistency matters in legal work because downstream readers assume the system treats like cases alike. When it does not, every output becomes suspect, even the correct ones. A system that is right in five different ways is harder to trust than one that is right the same way every time.
No. We grade AI outputs against the rubric and sources you provide. We are not lawyers, we do not employ licensed attorneys as graders, and our reports are not legal advice and do not replace review by qualified counsel. What you get is a documented, systematic read on where your system's outputs break your own standards.
Against the source materials you provide: your case database, your document set, your research corpus. We check that each citation exists, that it says what the output claims it says, and that it supports the point it is attached to. A real case cited for the wrong holding still fails.
Research memos, brief drafts, contract markup suggestions, clause summaries, deposition summaries, and similar text outputs. If your team can define what a correct output looks like and provide the sources, we can grade against them.
Yes, when you tell us the jurisdiction. The rubric includes a jurisdiction check: the output's rules, procedures, and citations must match the matter's actual venue. Mixed-jurisdiction answers are one of the most common failures we see in legal AI.
Attorney review is essential and expensive. We give you the systematic layer underneath it: every sample graded the same way, citation problems counted, failure patterns named. That lets your attorneys spend their time on judgment calls instead of catching the same fabricated citation repeatedly.
Hallucination detection asks whether outputs contain invented facts. Legal AI evaluation goes further: it checks whether real sources are used correctly, whether authority is overstated, and whether the output stayed in its approved role. For legal work, correct use of real sources matters as much as catching inventions.
Send 20 to 50 outputs with your sources. We will grade them all.
Not sure where to start? Talk to us and we will point you at the right evaluation.
Lock a quick 15-minute intro call — we'll scope your evaluation needs and deploy vetted experts within 48 hours.