A benchmark that measures what you ship

We design custom benchmarks built from your product's real tasks, so your scores track what users experience instead of what a public leaderboard rewards.

We design benchmarks and grading rubrics. We do not train models or sell rankings.

The short version

What is custom benchmark design?

Custom benchmark design is the work of building an evaluation suite around your product's actual tasks, data, and quality bar. Public benchmarks are generic by design; they cannot tell you whether your support bot resolves tickets or your code assistant writes working code. A good custom benchmark uses real task distributions, clear scoring rules, and versioning, so scores stay comparable from one release to the next and decisions get made on evidence.

Where it breaks

What goes wrong without your own benchmark

Teams without an internal standard argue about quality with anecdotes. A benchmark replaces anecdotes with numbers.

Leaderboard chasing

Optimizing for the wrong test

Public benchmark scores go up while user complaints stay flat, because the public test never measured your product. A custom benchmark scores the tasks your users actually perform, so improvements show up where they matter.

No internal standard

Every team grades differently

Product says quality is fine, support says it is broken, and both are working from vibes. A shared benchmark gives every team one score to argue about, with the task set and rubric visible to everyone.

Release decisions without data

Ship it and see

Model swaps and prompt changes go out on gut feel, and regressions get discovered by users. A versioned benchmark turns each release into a measured decision: better, worse, or unchanged, with the receipts.

Benchmark rot

Tests go stale

Static test sets get gamed, memorized, or quietly outdated as the product moves on. We design benchmarks for refresh and versioning: new tasks rotate in, old scores stay comparable, and the suite keeps measuring what you ship today.

How it works

From samples to answers in three steps

01

Share your product

You walk us through your product's real tasks and what good output looks like, or send task examples.

02

We design

We build the task set, the rubric, and the scoring rules, and you approve every part of it.

03

You get the benchmark

The full suite with baselines, versioning plan, documentation, and a walkthrough call.

Deliverables

What you get

  • Task set. A set of tasks sampled from your product's real distribution, documented and versioned.
  • Scoring rubric. Clear grading rules on the P0 to P3 scale, with worked examples for each level.
  • Baseline scores. Your current model graded on the new benchmark, so future releases have something to beat.
  • Versioning plan. How the benchmark refreshes over time without breaking score comparability.
  • Walkthrough call. We go through the benchmark, the baselines, and how to run it with your team.
20-50
pilot samples graded per batch
2-3
business days, once scope is confirmed
4
severity levels on every sample
Human-checked
reports
Why it matters

Why build your own benchmark

Your product is the test set.

Public benchmarks measure general capability across generic tasks. Your users do not ask generic questions; they bring your product's specific tasks, your domain's edge cases, and your quality expectations. Only a benchmark built from those can tell you whether a release made things better.

We sample the task set from your product's real distribution: the requests users actually make, weighted the way they actually occur. Scores from that benchmark predict user experience in a way no public leaderboard can.

Versioning keeps scores honest.

A benchmark is a measuring instrument, and instruments need calibration records. When tasks rotate, guidelines change, or the product moves on, scores from different versions stop being comparable unless the changes are tracked.

Our versioning plan covers exactly this: what changes between versions, how to compare scores across them, and when a refresh is due. You get a number you can trust in a meeting six months from now, not just today.

One standard for every team.

Product, engineering, support, and leadership usually carry four different opinions about quality, all argued from anecdotes. A shared benchmark replaces the anecdotes with one visible standard: the task set, the rubric, and the scores, open to everyone.

That shared standard changes how decisions get made. Model swaps become measured comparisons. Prompt changes get tested before they ship. Quality arguments end with evidence instead of whoever talks longest.

The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

Three things: a task set drawn from your product's real work, a rubric that defines good and bad output, and scoring rules that make results comparable over time. We build all three with you and document them so the benchmark outlives the engagement.

Longer than a grading pilot, because the task set and rubric need your input and approval. We start with a scoped design sprint on one task family, prove the approach, then expand. You approve every part before it is final.

We grade the baseline run as part of the design, so you see the benchmark working on your data. After handoff the benchmark is yours to run; we can also grade future releases as a separate engagement if you want an outside eye on the scores.

Yes, and for subjective quality it should. We design the human grading protocol alongside the automated checks: which samples get human review, how disagreements are settled, and what the agreement bar is.

Often enough to stay honest, rarely enough to keep scores comparable. Our versioning plan covers exactly this: new tasks rotate in on a schedule, version numbers mark the changes, and we document how to compare scores across versions.

Public benchmarks are useful for model selection, but they measure general capability, not your product. Our LLM evaluation service can show you the gap: grade your outputs on both and see which one predicts user complaints.

Yes, and that is one of its main uses. Grade outputs from each candidate model on the same task set with the same rubric, and you get a direct comparison on your product's work instead of generic leaderboard scores. Teams use this for model selection, vendor comparisons, and build-versus-buy decisions.

Measure what you actually ship.

Tell us about your product and get a scoped benchmark design plan.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.