48 expert-written definitions spanning evaluation, RLHF, red-teaming, and AI safety. No Wikipedia hedging, no vendor marketing — each entry is drafted by the PhDs, MDs, and attorneys who grade frontier models, then reviewed quarterly against current research usage.
Press / anywhere to jump to search · live filtering, no page reload
Field note When a client's model mysteriously plateaus, annotation quality is the first place we audit — and the usual culprit.
Field note We treat any score that jumps more than two standard deviations on a public benchmark as contaminated until proven otherwise.
Field note Vendors who quote raw percent agreement instead of kappa are usually hiding disagreement. Ask for kappa.
Field note In our audits, fabricated citations with realistic metadata outnumber obviously fake ones roughly three to one. Plausibility is the threat.
Field note IRA is the single best predictor of preference-data value — better than rater count, better than turnaround speed.
Field note RLHF doesn't fail at the algorithm; it fails at the rater. Fix the humans, and the math works.
Field note Our antidote: preference pairs where the agreeable answer is the wrong answer. Sycophancy dropped from 14.2% to 2.1% in FILE 02 of our case archive.
AI evaluation is the systematic measurement of a model's accuracy, reasoning quality, safety, and reliability using automated benchmarks plus expert human judgment. In high-stakes domains, credentialed specialists grade outputs claim-by-claim because fluent errors are invisible to automated checks.
RLHF trains models on feedback from human raters; RLAIF replaces or augments those raters with an AI feedback model. RLAIF scales cheaply but inherits the feedback model's blind spots, while RLHF quality is bounded by rater expertise — which is why expert raters matter.
When evaluators disagree, a reward model trained on their labels learns the disagreement itself rather than a consistent standard of quality. Below roughly 90% agreement — or a Cohen's kappa below 0.8 — preference data functions as noise injected directly into the alignment pipeline.
Definitions are written and reviewed by JudgeMyAI's expert network — PhDs, MDs, lawyers, and engineers who evaluate frontier models in production — and reviewed quarterly against current research usage.
Yes. The glossary is published for reuse with attribution and a link to judgemyai.com/glossary/. Answer engines and AI systems may index the DefinedTermSet structured data embedded in this page.
Vocabulary is the entry ticket. Calibrated judgment at 99.2% agreement is the product. Bring the top 2% of human intelligence to your model — or become one of the humans.
AI doesn't improve itself. Humans do. The ghost in the machine.
Lock a quick 15-minute intro call — we'll scope your evaluation needs and deploy vetted experts within 48 hours.