Pick the right model with evidence.

You are choosing between two or three models and the demos all look fine. We compare LLM models on your own tasks with your own rubric, so the decision comes from graded samples instead of vibes.

We grade outputs. We compare models on evidence from your use case. We do not train models, run fine-tuning, or sell experts we cannot verify.

20 to 50
samples graded per model
2 to 3
business days, once scope is confirmed
4
severity levels, P0 to P3
Human-checked
reports
The short version

What is model comparison?

Model comparison is a side by side evaluation of two or more language models on the same set of tasks, graded against the same rubric. You need it when you are choosing a model for production, switching providers, or asking whether a smaller, cheaper model is good enough for your workload. Good looks like graded samples for every candidate, identical grading rules for all of them, and a recommendation you can defend in a meeting. It is a close cousin of LLM evaluation, pointed at one question: which model earns a place in production.

Where it breaks

Demos hide the differences

Models feel identical on easy questions. The gaps show up on your real tasks, late, and in production.

Cherry-picked demos

The demo that sold you

Every vendor demo shows the model at its best. Your users ask worse questions with messier context. Graded samples on your real prompts show what the demo never did, including the failures you would have met in week one of production.

Benchmark worship

Public benchmarks lie about your use case

A model can top a public leaderboard and still fail your onboarding flow. Benchmarks test general knowledge. Your product needs task-specific grading on the exact work the model will do, which is what a custom benchmark design gives you.

Cost blind spots

The cheaper model that costs more

A smaller model looks like a saving until error rates climb and support tickets follow. We grade output quality per model so you can weigh real quality against real price, per task, before you commit to a switch.

Silent switching

The provider changes the model under you

Model versions update, APIs drift, and the model you tested in March is not the model serving in June. A graded baseline from a comparison pilot lets you spot when a switch, yours or theirs, changed your quality.

The report

What the comparison shows, per model

One task set, one rubric, every candidate graded the same way.

DimensionWhat you get
Sample setThe same 20 to 50 tasks run through each model, so the comparison is fair from the start.
GradingEvery sample graded P0 to P3 by human reviewers against one shared rubric, with a written reason.
Severity mixWhere each model's P0 and P1 failures cluster, broken down task by task.
Failure patternsThe recurring mistakes each candidate makes on your tasks, described in plain words.
RecommendationWhich model fits your use case, what it costs you in quality, and what to re-test before launch.
How it works

From candidates to answers in three steps

01

Send the candidates

Tell us which two or three models you are deciding between and share 20 to 50 real tasks. You can send outputs you already generated, or send prompts and we will generate the samples.

02

We grade

Human reviewers grade every sample against one shared rubric, the same rules for every model, backed by automated checks. No candidate gets a home-field advantage.

03

You get the report

A side by side report with grades per model, failure patterns, severity mix, and a recommendation you can take to your team. Plus a walkthrough call.

Deliverables

What you get

  • Per-model grade sheetsEvery sample from every candidate scored P0 to P3 with a written reason.
  • Severity mix comparisonWhere each model's P0 and P1 failures cluster, task by task.
  • Failure pattern mapThe recurring mistakes each candidate makes on your tasks, described plainly.
  • Selection recommendationA defended recommendation for your use case, plus what to re-test before launch.
Reading the results

What to look for in the results

Start with the severity mix, not the average

When you open the report, skip the averages and go straight to the P0 and P1 counts per model. Two models can share the same average score while failing in completely different ways: one invents facts rarely but catastrophically, the other makes small errors constantly. The severity mix shows you which kind of wrong you are buying. Our severity scale explains what each level means in practice.

Then read the failure patterns

The pattern map is where the decision gets made. If Model A fails on your hardest task type and Model B fails on trivia you can guardrail away, the choice is clear even when the totals are close. Bring the pattern map to your team meeting. It answers the question "what breaks" better than any single number, and it tells your engineers exactly what to fix first after you choose.

One more thing the report gives you: a re-test list. Once you pick a model, its failure patterns become your regression set. Run the same graded samples after every prompt change, every retrieval tweak, and every provider version bump, and you will know immediately whether a change helped or hurt. That is how a one-time comparison turns into an ongoing quality habit with no extra setup. Teams that skip this step end up re-deciding the same model question every six months. Teams that keep the graded baseline answer it once, with evidence, and move on to building.

The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

Two or three is the sweet spot. Beyond that, the sample set gets thin per model and the comparison loses its bite. If you are choosing between four or more, talk to us first and we will suggest a shortlist round.

Either way works. You can send outputs you already generated, or you can send prompts and task descriptions and we will generate the samples for grading. The rubric can be yours, or we can help you write one before grading starts.

Yes. Open-weight models against hosted APIs, large against small, your current model against a candidate replacement. The grading stays identical across candidates, so the comparison is fair.

One shared rubric, the same task set, human reviewers grading every sample, and a severity with a written reason on each one. Nothing is averaged away or hidden behind a single score. If you want the full method, it is in our methodology.

Then we say so, and we show you where they differ anyway. One model may fail loudly and rarely while another fails quietly and often. That difference matters in production, and it is in the report either way.

A standard pilot grades one model deeply on your tasks. A comparison pilot grades two or three candidates on the same task set and answers one question: which one do we ship. If you already picked a model, start with the pilot.

Stop guessing which model wins.

Send the candidates. We will show you the grades.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.