Preference data your reward model can trust

We produce preference data RLHF teams can trust: labeled preference pairs for your training, with human-chosen winners and losers and written reasons for every choice.

We produce labeled preference pairs for your training. We do not run training, fine-tuning, DPO, or RLHF loops ourselves.

The short version

What is preference data for RLHF?

Preference data for RLHF is a dataset of prompts with pairs of responses, where a human has marked which response is better and explained why. Reward models learn from those choices, so the quality of the labels sets the ceiling for the whole effort: noisy labels teach noisy preferences. Good preference data comes from clear guidelines, trained labelers, and a documented reason behind every choice. We produce the dataset. Your team runs the training.

Where it breaks

Bad preference data teaches bad preferences

A reward model is only as good as the labels behind it. Here is where preference datasets go wrong.

Noisy labels

Labelers guessing

Vague guidelines produce random preferences, and a reward model trained on random preferences learns nothing useful. We write tight labeling guidelines with examples, train labelers on them, and check agreement before full production.

Forced winners

Both answers are fine

Forcing a winner on near-identical responses adds pure noise to the dataset. We allow ties and record confidence, so your training pipeline can weight or filter the close calls instead of learning from coin flips.

Guideline drift

Labelers invent their own rules

Over a long labeling run, standards slip: labelers start applying personal taste instead of the guideline. We run calibration rounds and ongoing agreement checks, and the report shows you the numbers.

Missing reasons

A label with no explanation is a dead end

When a preference looks odd six months later, the written reason is the only way to understand it. Every pair we ship carries the labeler's reasoning, so your team can audit, debug, and refine the dataset.

How it works

From samples to answers in three steps

01

Send prompts

You send the prompts to label, with candidate responses, or we generate response pairs from your models.

02

We label

Trained labelers pick winners and losers against your guidelines and write the reason for each choice.

03

You get the dataset

Labeled pairs in your format, with agreement stats, guideline notes, and a walkthrough call.

Deliverables

What you get

  • Labeled preference pairs. Chosen and rejected responses per prompt, with ties and confidence marked, in the format you specify.
  • Written reasons. The reasoning behind every preference choice, so the dataset stays auditable.
  • Labeling guidelines. The exact instructions our labelers worked from, yours to keep and reuse.
  • Agreement report. Inter-labeler agreement numbers and how disagreements were resolved.
  • Walkthrough call. We go through the dataset and the labeling decisions with your team.
20-50
pilot samples graded per batch
2-3
business days to your report
4
severity levels on every sample
100%
of reports reviewed by a human
Why it matters

Why label quality decides RLHF outcomes

Reward models inherit your labels.

A reward model learns one thing: which responses your labelers preferred. Every inconsistency, shortcut, and guess in the labeling becomes part of what the model optimizes for. There is no step later in the pipeline that fixes bad labels; the reward model amplifies them.

This is why preference data deserves the same rigor as any training input your team would never ship unchecked. Clear guidelines, trained labelers, agreement measurement, and written reasons are not overhead. They are the product.

Guidelines are the real product.

Ask five people what a good response looks like without a guideline and you get five philosophies. Preference labeling without tight guidelines produces noise, and noise in, noise out. The guideline is what turns individual taste into a consistent standard.

We draft guidelines with worked examples, calibrate labelers on them, and measure agreement before production labeling. You approve the final version and keep it, so future labeling rounds and internal teams work from the same standard.

Reasons make datasets auditable.

Six months from now, someone will look at a preference pair and ask why the labeler chose the way they did. Without a written reason, that question has no answer, and the dataset slowly becomes unmaintainable.

Every pair we ship carries the labeler's reasoning. That makes the dataset debuggable: you can find the guideline that produced odd labels, fix the guideline, and know which pairs to revisit. Datasets with reasons age well; datasets without them do not.

The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

No, plainly. We produce labeled preference pairs for your training. We do not run training, fine-tuning, DPO, or RLHF loops. If you need training run, that is your team's work or a different vendor's; our lane is the data.

Whatever your pipeline expects: JSONL with prompt, chosen, rejected, and reason fields is the common shape, but we match your schema. Tell us the format up front and the first delivery fits it.

It depends on what you are training and how noisy the task is. A pilot batch lets us measure labeler agreement and guideline clarity first; then we scope the full volume with real numbers instead of guesses.

Disagreement is signal, not failure. We measure it, resolve the clear cases with a senior reviewer, and surface the genuinely ambiguous ones to you with both arguments. The agreement report shows exactly how much disagreement there was and where.

Yes. Send the prompts with your models' candidate responses and we label preferences over them. Or we can generate the pairs from models you point us at. Either way, the labeling standard is your guideline, approved by you.

Preference labeling is one specialized kind of annotation: ranking outputs by quality against a guideline. Our data annotation service covers the broader set: categories, spans, scores, and other label types.

Yes. Send them and we calibrate our labelers to your guideline, then report back where it needed clarification or extra examples. If you do not have guidelines yet, we draft them with you: you bring the domain knowledge, we bring the structure that makes guidelines labelable. You approve the final version either way.

Get preference labels you can defend.

Send your prompts and get a labeled pilot batch with agreement stats.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.