The AI severity scale JudgeMyAI uses on every sample: P0 Critical, P1 Major, P2 Minor, P3 Clean. One shared language for how wrong an output is, so engineers, reviewers, and founders all mean the same thing when they say "this failed."
We grade outputs. We do not train models, run fine-tuning, or sell experts we cannot verify.
The P0 to P3 severity scale is a four-level grading system for AI outputs. P0 means critical: the sample fails outright. P1 means major: it needs rework before it can ship. P2 means minor: a small deduction for a polish issue. P3 means clean: accurate, complete, well formed. Every sample we grade gets one of these four levels, a written reason, and a suggested fix. It is the backbone of our methodology.
Without a shared scale, every quality conversation turns into an argument about words.
How bad? Bad enough to block the launch, or bad enough to fix next sprint? Without levels, "bad" covers everything from a typo to a fabricated medical claim, and the team cannot prioritize. Severity levels turn adjectives into decisions.
Engineering says "sev 1," support says "urgent," the founder says "this is fine." Three vocabularies, zero shared meaning. One scale across the company means a P0 in a report means the same thing to everyone who reads it.
An average of 4.2 out of 5 looks healthy until you learn it contains three P0s. Averages are where critical failures go to hide. A severity scale keeps every failure visible at its own level, because nothing gets averaged away.
"This output is a 2." A 2 of what? What do I change? A severity level plus a written reason tells the prompt engineer exactly what broke and how bad it was, which is the minimum useful unit of grading feedback.
Here is the same scenario at each severity level, so you can feel the difference. The setup: a customer asks a support chatbot, "What is your refund policy?" The real policy, from the company's help center: full refund within 14 days of purchase, no questions asked.
"You can get a full refund within 14 days of purchase, no questions asked. Just head to Settings, then Billing, and click Request refund." Every claim matches the source. The steps are right. This is the bar: accurate, complete, well formed.
"You can get a full refund within 14 days of purchase, no questions asked! Just head to Settings, then Billing, and click Request refund. Our team usually processes these within a couple of days." The policy is right and the steps are right. "A couple of days" is slightly vague against a documented one-business-day target, but it does not change what the customer does. Small deduction, not a trust issue.
"You can get a full refund within 30 days of purchase. Just contact support and they will sort it out." The window is wrong, doubled, and the self-serve path is missing. A customer reading this will wait too long or take the slower route, and either way the company eats the confusion. Materially misleading. Cannot ship as is.
"You are entitled to a full refund within 90 days under our lifetime satisfaction guarantee, and if we refuse, you can file a chargeback which we are legally required to honor." Nothing in this is true. The policy is invented, the "lifetime satisfaction guarantee" does not exist, and the legal claim is fabricated with total confidence. A customer acting on this creates a dispute the company cannot win gracefully. This fails the sample outright.
Notice what the scale asks of the grader: which bucket does this fall in? Not "how many points out of a hundred," not "rate from 1 to 5." Four buckets, each with a clear decision attached: fails, rework, small fix, clean. Graders pick consistently because the boundaries are about consequences, and consequences are easy to judge. That is the whole design philosophy: grade the impact, not the vibe. If you are writing your own evaluation rubric, steal this structure.
You send 20 to 50 model outputs plus your rubric, or we help you write one.
Trained reviewers grade every sample against the rubric, backed by automated checks.
Sample-level grades, issue patterns, recommended fixes, and a walkthrough call.
Because graders cannot reliably tell fine-grained levels apart. With ten levels, the boundary between 6 and 7 is a coin flip, and that noise pollutes your results. Four levels map to four real decisions: fails, needs rework, minor fix, clean. Every level earns its place.
The sample gets the worst level, and the report lists every issue found. A P0 plus a P2 is still a P0 sample, because the critical failure decides whether it ships. The lesser issues are still recorded so nothing is lost.
Yes. The four levels stay the same everywhere; what changes is the rubric that defines what counts as P0 versus P1 for your product. A wrong refund window is a P1 for a store and might be a P0 for a bank. The scale is shared, the calibration is yours.
Please do. It works just as well for internal QA: tag every reviewed output P0 to P3, keep the reasons, and watch your P0 rate over time. If you want help calibrating your team on it, that is part of what a pilot covers.
Automated judges can assign severity levels too, and we use them for scale. But the level definitions, the examples, and the edge-case rulings come from human judgment first. Automation inherits the scale; it does not invent it. Our automated evaluation page goes deeper.
In our sample report: real graded samples with severity levels, reasons, and suggested fixes. It is the fastest way to see what P0 through P3 look like on actual outputs.
Send 20 to 50 outputs. Get severity levels, reasons, and fixes in 2 to 3 business days once scope is confirmed.
Not sure where to start? Talk to us and we will point you at the right evaluation.
Lock a quick 15-minute intro call — we'll scope your evaluation needs and deploy vetted experts within 48 hours.