CSAT up 12 points after expert-driven RLHF on a support agent
ChallengeA Series-C SaaS company's AI support agent was confidently inventing refund policies and feature promises — 8.9% of policy answers were wrong in ways that created real obligations. Customers noticed; churn interviews cited "the bot lied to me."
Engagement120 product-expert evaluators, certified on the client's actual knowledge base, produced 500k preference pairs rewarding verified-policy fidelity over fluent improvisation, plus a live escalation rubric that routes uncertain answers to humans.
- Delivered: 500k preference pairs plus a live escalation rubric for uncertain answers.
- Cohort: 120 evaluators certified on the client's knowledge base.
- Status: Continuous program renewed three times.