Wrong answers teach wrong lessons

EdTech AI evaluation that grades your tutoring, quiz generation, and explanation outputs against your curriculum rubric. A wrong explanation inside a product that teaches is not a minor bug. It is the lesson the student walks away with.

We grade outputs. We do not train models, run fine-tuning, or sell experts we cannot verify.

20-50
pilot samples graded
2-3
business days turnaround
4
severity levels per sample
100%
human-reviewed reports
The short version

What is EdTech AI evaluation?

EdTech AI evaluation is grading what a learning product's model tells students: explanations, hints, quiz questions, feedback on answers, and full tutoring conversations. It is needed by anyone shipping AI into classrooms, homework apps, or test prep. Good means correct content, teaching that builds understanding instead of handing out answers, and language right for the student's level.

Where it breaks

The failures EdTech AI evaluation catches

Students trust the tutor. When the tutor is wrong, nobody in the room knows.

Failure pattern

Plausible wrong explanations

The model explains a math step or a science concept with total confidence and one wrong move in the middle. The explanation reads smoothly, so the student memorizes the wrong method. These are the highest-stakes hallucinations in education: wrong, teachable, and invisible to the learner.

Failure pattern

Answers instead of teaching

A homework helper that hands over the final answer on the first message is not tutoring. Good pedagogy scaffolds: hints, questions back, smaller steps. We grade whether the model actually teaches or just completes the assignment for the student.

Failure pattern

Wrong level for the student

A fifth-grade explanation written with college vocabulary. A quiz for beginners that assumes the advanced unit. Level mismatch does not look like an error, but it ends learning just as surely. We grade against the age band and proficiency level you specify.

Failure pattern

Slips in examples and tone

Word problems that lean on tired stereotypes. Feedback that shames a wrong answer instead of correcting it. Encouragement that tips into empty praise. These are P1 and P2 issues individually, but across thousands of students they shape how kids see themselves.

What we actually check

Three things, on every sample

Correctness of the content

Every factual claim, worked step, and quiz answer gets checked against your curriculum sources. In multi-step reasoning we grade the steps, not just the final answer, because a right answer reached by a wrong method still teaches the wrong method. For generated quizzes we also check that the keyed answer is actually correct and the distractors are plausible but wrong, since a bad distractor gives the answer away. Send your answer keys or reference material and we build them into the rubric.

Pedagogical soundness

Does the model guide or just give? We grade hint quality, the use of questions that prompt thinking, and whether feedback on a wrong answer identifies the actual misconception. This is where a well-written evaluation rubric matters: teaching quality has to be defined before it can be graded.

Age and level fit

Vocabulary, sentence complexity, assumed background knowledge, and the emotional tone of feedback. A sample can be fully correct and still fail if it talks past the student it was written for. We grade fit against the specific band you ship to.

Encouragement and tone

How the model talks to a student who is wrong matters as much as the correction itself. Feedback that shames, praise that is empty, and encouragement that never connects to the actual work all get graded. We look for tone that keeps the student trying: specific about what was right, clear about what to fix, and honest instead of flattering. Across a batch, this reveals whether your tutor builds confidence or just performs friendliness.

How it works

From samples to answers in three steps

01

Send samples

You send 20 to 50 model outputs plus your rubric, or we help you write one. Include answer keys and curriculum standards if you have them.

02

We grade

Trained reviewers grade every sample against the rubric, backed by automated checks for consistency and coverage.

03

You get the report

Sample-level grades, issue patterns, recommended fixes, and a walkthrough call.

Deliverables

What you get

  • Sample-level gradesEvery sample scored P0 to P3 with a written reason, including which step of a worked solution failed.
  • Issue summaryThe patterns across the batch, ranked by severity and frequency, mapped to subjects and levels.
  • Recommended fixesConcrete next steps for your prompts, your content guidelines, or your guardrails.
  • Walkthrough callWe go through the report with your team and answer questions.
The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

Yes. Send the standards, scope and sequence, or textbook chapters the model should follow, and we write them into the rubric. A sample that is correct but off-curriculum gets flagged, because in a classroom, the wrong chapter is the wrong answer.

Yes, and we grade the steps, not just the final answer. A correct final answer with a broken step in the middle is graded as a P1, because the student learns the steps. Send your worked answer keys and we check against them.

Anything unsafe for the age band, from self-harm content to age-inappropriate material, is graded P0 and surfaced first in the report. Safety is always the top line of an EdTech report, before any discussion of quality.

No. We grade model outputs against your rubric; we do not run classroom studies or user research. If you have transcripts of real student sessions you want graded, we can review those as samples.

Yes. Tell us the subjects and levels when you send samples, and we spread the batch across them with subject-specific criteria in the rubric. Each subject gets its own pattern summary in the report.

Start a pilot: 20 to 50 samples, your rubric or ours, graded in 2 to 3 business days once scope is confirmed with a walkthrough call. That first report usually tells a team exactly where their tutor stands.

Know what your tutor is actually teaching

Send 20 to 50 samples. Get every one graded against your curriculum.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.