A pilot that answers one question

How to run an AI evaluation pilot: scope it to one question, grade 20 to 50 samples against a written rubric, and walk away knowing what to do next. A pilot is the smallest evaluation that still tells the truth.

We grade outputs. We do not train models, run fine-tuning, or sell experts we cannot verify.

The short version

What is an AI evaluation pilot?

An AI evaluation pilot is a small, time-boxed evaluation: 20 to 50 model outputs graded against a written rubric, delivered with a report in 2 to 3 business days once scope is confirmed. It exists to answer one question, usually "is this ready?" or "what is actually wrong?" A pilot is not a full audit and not a monitoring program. It is the first real measurement, sized so the results arrive while they are still useful.

Where it breaks

How pilots go wrong

A pilot is small, which makes it easy to waste. These are the usual ways.

Failure pattern

The pilot with five questions

"Is it ready, which model is better, what should we fix, and how do we compare to competitors?" One pilot cannot answer five questions with 40 samples. Scope creep turns a sharp pilot into a blurry one. Pick one question and let the pilot answer it properly.

Failure pattern

Samples picked to look good

The team sends its best outputs, consciously or not, and the pilot confirms what they already believed. A pilot that cannot surprise you is theater. Include the hard cases, the edge cases, and a few outputs you already suspect are bad.

Failure pattern

No rubric, just vibes

Samples go out with "tell us what you think" instead of a written standard. The graders improvise, the results argue with themselves, and the report reads like five opinions. Write the rubric before the pilot starts, even a rough one.

Failure pattern

Results nobody acts on

The report lands, everyone nods, and nothing changes. A pilot needs an owner and a decision waiting on it: a launch date, a prompt rewrite, a model choice. No pending decision, no pilot. That rule alone saves most wasted evaluations.

Running it

Pick one question

Everything in a pilot flows from the question. Write it down in one sentence before you do anything else. Good pilot questions sound like this:

Questions that work

"Is our support chatbot's answer quality good enough to launch?" "What are the top three failure patterns in our RAG answers?" "Did the new prompt actually improve factual accuracy?" Each one names the product, the dimension, and the decision waiting on the answer. Our LLM evaluation guide covers how this question shapes everything downstream.

Questions that do not

"How good is our AI?" Too broad to grade. "Is model A better than model B at everything?" Too broad and too comparative for 40 samples. "Can you check our outputs?" Not a question at all. If your question needs the word "everything," narrow it until it names one thing.

Let the question size the batch

A launch-readiness question needs coverage across your real input mix, so go toward 50 samples with deliberate strata. A "did the prompt improve" question needs before-and-after pairs on the same inputs, so 20 to 30 pairs beat 50 random singles. Our sampling guide walks through the details.

Scope check

What a pilot is not

Pilots get oversold. Here is what a 20-to-50-sample pilot cannot do, so you scope honestly.

Not a full audit

A pilot finds the repeating patterns; it does not certify the product. If you need comprehensive coverage across every task type, language, and edge case, that is a full evaluation audit, a bigger engagement with a bigger batch. The pilot often tells you whether the audit is worth doing.

Not a benchmark

A pilot grades your outputs against your rubric. It does not rank you against competitors or produce a leaderboard number. If you need model comparison, say so up front: comparing two models doubles the batch and changes the sampling, and that is a different pilot design.

Not training data

The graded samples are measurements, not a dataset for fine-tuning. We do not train models, and a pilot's 40 samples would not train one anyway. If you need labeled data to teach a model, that is preference data work, a separate engagement with different sampling and labeling.

How it works

From samples to answers in three steps

01

Send samples

You send 20 to 50 model outputs plus your rubric, or we help you write one. We confirm the question and the batch before grading.

02

We grade

Trained reviewers grade every sample against the rubric, backed by automated checks.

03

You get the report

Sample-level grades, issue patterns, recommended fixes, and a walkthrough call where we answer your one question directly.

Deliverables

What you get

  • Sample-level gradesEvery sample scored P0 to P3 with a written reason.
  • Issue summaryThe patterns across the batch, ranked by severity and frequency, tied back to your one question.
  • Recommended fixesConcrete next steps for your prompts, retrieval, or guardrails.
  • Walkthrough callWe go through the report with your team and answer questions.
The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

Three things: your one question, 20 to 50 sample outputs, and your rubric or a draft of one. If you have fact sources, a style guide, or answer keys, send those too. We confirm scope and the batch with you before any grading begins.

Two to three business days from the moment scope and samples are confirmed, including the walkthrough call. If your rubric needs writing from scratch, add the time we spend drafting it with you, which usually happens in the same window.

Every sample with its P0 to P3 grade and written reason, the issue patterns ranked by severity and frequency, recommended fixes, and a direct answer to your one question. See our sample report for the format.

Three common paths: the pilot answers the question and you act on it; you fix what it found and run a second pilot to confirm; or the findings justify a full audit or continuous monitoring. The walkthrough call covers which path fits your results.

Yes, with a different design: both models answer the same inputs, and we grade the pairs blind. Say you want a comparison when you reach out, because it changes the sampling and the batch size. Our model comparison page covers the full version.

If you have one open question about your model's outputs and 20 to 50 samples to send, yes. If you are pre-launch with no outputs yet, generate a realistic batch first. If you already know you need comprehensive coverage, ask us about a full audit instead. Talk to us and we will tell you honestly which fits.

Answer your one question

Send 20 to 50 samples. Get the answer, with the evidence, in 2 to 3 business days once scope is confirmed.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.