Case Studies: AI Evaluation, RLHF & Red-Teaming Results | JudgeMyAI
Case Studies · Declassified Archive

Engagement files from the alignment frontier — names sealed, numbers real.

Below are 22 declassified case files from 214 completed engagements. Client identities stay sealed under NDA — REDACTED REDACTED — but the metrics don't: every figure was measured against client-side analytics before and after our work. Filter by sector, open any file, read the full record.

NDA policy: we never name clients, model codenames, or proprietary benchmarks. Reference calls with matching-industry clients are available under mutual NDA.
31.4MExpert evaluations delivered
12,480Vetted experts in talent graph
99.2%Avg. inter-rater agreement
−87%Median hallucination reduction
9 daysAvg. time to calibrated cohort
FILE 16 // JMA-2026-0163 RLHF Program 2026 · 16 weeks · 120 experts

CSAT up 12 points after expert-driven RLHF on a support agent

ChallengeA Series-C SaaS company's AI support agent was confidently inventing refund policies and feature promises — 8.9% of policy answers were wrong in ways that created real obligations. Customers noticed; churn interviews cited "the bot lied to me."

Engagement120 product-expert evaluators, certified on the client's actual knowledge base, produced 500k preference pairs rewarding verified-policy fidelity over fluent improvisation, plus a live escalation rubric that routes uncertain answers to humans.

Hallucinated policy claims −94% CSAT +12 pts in two quarters Escalation precision 96%
Duration16 wks
  • Delivered: 500k preference pairs plus a live escalation rubric for uncertain answers.
  • Cohort: 120 evaluators certified on the client's knowledge base.
  • Status: Continuous program renewed three times.
FILE 17 // JMA-2025-0951 Benefits QA 2025 · 7 weeks · 63 experts

HR-benefits copilot stops giving wrong plan advice

ChallengeAn HR-tech platform's benefits copilot recommended wrong plans 7.3% of the time — advising an employee to pick a plan that would cost them thousands. The failure mode: plans differ in dense, boring ways that generic raters skim past.

Engagement63 benefits-certified (CEBS) evaluators graded 90k advice conversations against plan documents, treating every dollar-quantitative claim as verify-or-flag. Error taxonomy drove both fine-tuning and a hard calculator integration for cost comparisons.

Wrong-plan advice 7.3% → 0.6% 90k conversations audited Open-enrollment incidents: zero
Duration7 wks
  • Delivered: 90k audited advice conversations plus a verify-or-flag claim rubric.
  • Cohort: 63 CEBS-certified benefits evaluators.
  • Status: Zero open-enrollment incidents post-launch.
FILE 18 // JMA-2026-0055 Extraction Eval 2026 · 6 weeks · 88 experts

98.7% attribute accuracy for an e-commerce catalog pipeline

ChallengeA marketplace processing 4M product listings needed spec extraction (wattage, materials, compatibility) accurate enough to power search filters. Vendor raters plateaued at 94% — each miss a wrong search result or a return.

Engagement88 engineers and category specialists built a graded hierarchy: unambiguous specs, judgment-call specs, and unverifiable claims — each with its own verification rule. The labeled set retrained extraction and defined when the pipeline must abstain instead of guess.

Attribute accuracy 94% → 98.7% Bad-spec returns −34% Silent errors cut 6×
Duration6 wks
  • Delivered: Three-tier extraction verification hierarchy plus an abstention policy.
  • Cohort: 88 engineers and category specialists.
  • Status: Live across 4M listings today.
FILE 19 // JMA-2025-0689 Freshness Audit 2025 · 9 weeks · 70 experts

62k documents re-scored after an internal-knowledge assistant went stale

ChallengeA 9,000-person company's internal AI assistant answered policy questions from stale documents — org changes, pricing, security procedures — with total confidence. No one knew which answers were current.

Engagement70 evaluators with tenure inside the client's functions re-scored 62k documents for currency and authority, then graded the assistant's answers against the corrected corpus. Freshness signals now gate retrieval, and a standing panel re-audits quarterly.

Stale answers served −81% 62k documents freshness-scored Quarterly re-audit retained
Duration9 wks
  • Delivered: 62k freshness and authority scores feeding retrieval gating.
  • Cohort: 70 evaluators with direct tenure in the client's functions.
  • Status: Standing quarterly re-audit panel retained.

Open-Source Model Developers

Files 20–22 · Lab-grade alignment on community budgets
FILE 20 // JMA-2026-0071 Preference Set 2026 · 10 weeks · 150 experts

Lab-grade RLHF for a 70B open model

ChallengeAn open-source collective fine-tuning a 70B base model had compute and enthusiasm but no access to expert preference data. Community-sourced rankings were dominated by style preference — the model was getting eloquent, not correct.

Engagement150 experts produced a 240k-pair preference set under a subsidized open-license program, with rationales and calibration checks. The collective trained its reward model on the set and released everything — data included — under an open license.

240k expert pairs, open-released Expert-judged quality +0.84 / 10 pts Style-over-substance failures −71%
Duration10 wks
  • Delivered: 240k-pair open preference set with rationales and calibration checks.
  • Cohort: 150 experts under a subsidized open-license program.
  • Status: Data and model both open-released to the community.
FILE 21 // JMA-2025-0838 Safety Fine-Tune 2025 · 8 weeks · 47 experts

Jailbreak success rate 31% → 4% for a community model

ChallengeA widely-downloaded community model was trivially jailbreakable — 31% success in the maintainers' own testing, worse in the wild. Automated safety fine-tuning had plateaued, and the model was being used in moderation tooling.

Engagement47 red-team specialists ran a structured attack campaign, then built adversarial-preference data pairing successful jailbreaks with refusals that stay helpful on legitimate adjacent requests — refusing without becoming useless.

Jailbreak success 31% → 4% False-refusal rate under 2% Safety set open-released
Duration8 wks
  • Delivered: Adversarial-preference safety set balancing refusal with helpfulness.
  • Cohort: 47 red-team specialists running structured attack campaigns.
  • Status: Safety set open-released and community-adopted.
FILE 22 // JMA-2026-0125 Multilingual Eval 2026 · 12 weeks · 96 experts

Native-expert evaluation across 12 languages, including 4 low-resource

ChallengeA multilingual open model posted strong English benchmarks but had never been honestly evaluated in its target languages. Machine-translated test sets were measuring translation quality, not reasoning — and cultural-context failures were invisible.

Engagement96 native-speaker PhDs and professional linguists across 12 languages evaluated reasoning, safety, and cultural-context handling in-language, not in translation. The audit surfaced a toxicity-detection blind spot in two low-resource languages and a 2.3× quality gap that English benchmarks had completely hidden.

Cultural-context failures −63% 2.3× hidden quality gap closed 12 languages · in-language
Duration12 wks
  • Delivered: In-language evaluation across 12 languages plus a blind-spot report.
  • Cohort: 96 native-speaker PhDs and professional linguists.
  • Status: Fine-tune closed the 2.3× gap; blind spot patched.
Core Competencies

The capabilities behind every file above.

  • LLM Response EvaluationExpert-domain grading of large language model outputs for factual accuracy, reasoning quality, tone, and instruction adherence, performed by credentialed subject-matter specialists.
  • RLHF Preference Data CollectionHuman preference rankings, pairwise comparisons, and reward-model training data produced by top-2% experts for reinforcement learning from human feedback pipelines.
  • AI Red-TeamingSystematic adversarial testing of LLMs to identify jailbreaks, unsafe outputs, bias, and alignment failures before deployment.
  • Hallucination Detection & PreventionClaim-by-claim verification of model-generated facts, citations, and reasoning chains, with labeled failure data used for fine-tuning and guardrail construction.
  • AI Safety & Compliance EvaluationAssessment of model behavior against medical, legal, and financial safety standards, including regulated-industry documentation and audit support.
  • Domain-Specific Model TrainingCurriculum design and expert-led fine-tuning data for specialized fields including medicine, law, engineering, and scientific research.
Straight Answers

Frequently asked, honestly answered.

Every JudgeMyAI engagement is covered by NDA. Client identities, model codenames, and proprietary benchmarks are redacted. All metrics shown are verified against client-side analytics and reproduced here with written permission.

Yes. Every figure comes from client-verified measurement windows before and after engagement, with methodology documented in the delivered report. Where a client's internal benchmark is referenced, the metric is indexed rather than exposed.

Engagements range from 4-week red-team sprints to 40-week continuous evaluation programs, staffed by 15 to 340 domain experts depending on scope, with calibrated cohorts producing graded evaluations, preference data, or adversarial findings.

Frontier AI labs, healthcare and medical AI, finance AI, legal AI, enterprise SaaS, and open-source model developers — industries where hallucinations, fabricated citations, or safety failures carry material consequences.

Yes. Under mutual NDA, we arrange reference calls with existing clients in matching industries and engagement types.

Your model's case study starts with one decision.

Every file in this archive began with a team that refused to ship unverified intelligence. Bring the top 2% of human judgment to your model — or become one of the humans.

Apply for AI Jobs

AI doesn't improve itself. Humans do. The ghost in the machine.