EdTech AI evaluation that grades your tutoring, quiz generation, and explanation outputs against your curriculum rubric. A wrong explanation inside a product that teaches is not a minor bug. It is the lesson the student walks away with.
We grade outputs. We do not train models, run fine-tuning, or sell experts we cannot verify.
EdTech AI evaluation is grading what a learning product's model tells students: explanations, hints, quiz questions, feedback on answers, and full tutoring conversations. It is needed by anyone shipping AI into classrooms, homework apps, or test prep. Good means correct content, teaching that builds understanding instead of handing out answers, and language right for the student's level.
Students trust the tutor. When the tutor is wrong, nobody in the room knows.
The model explains a math step or a science concept with total confidence and one wrong move in the middle. The explanation reads smoothly, so the student memorizes the wrong method. These are the highest-stakes hallucinations in education: wrong, teachable, and invisible to the learner.
A homework helper that hands over the final answer on the first message is not tutoring. Good pedagogy scaffolds: hints, questions back, smaller steps. We grade whether the model actually teaches or just completes the assignment for the student.
A fifth-grade explanation written with college vocabulary. A quiz for beginners that assumes the advanced unit. Level mismatch does not look like an error, but it ends learning just as surely. We grade against the age band and proficiency level you specify.
Word problems that lean on tired stereotypes. Feedback that shames a wrong answer instead of correcting it. Encouragement that tips into empty praise. These are P1 and P2 issues individually, but across thousands of students they shape how kids see themselves.
Every factual claim, worked step, and quiz answer gets checked against your curriculum sources. In multi-step reasoning we grade the steps, not just the final answer, because a right answer reached by a wrong method still teaches the wrong method. For generated quizzes we also check that the keyed answer is actually correct and the distractors are plausible but wrong, since a bad distractor gives the answer away. Send your answer keys or reference material and we build them into the rubric.
Does the model guide or just give? We grade hint quality, the use of questions that prompt thinking, and whether feedback on a wrong answer identifies the actual misconception. This is where a well-written evaluation rubric matters: teaching quality has to be defined before it can be graded.
Vocabulary, sentence complexity, assumed background knowledge, and the emotional tone of feedback. A sample can be fully correct and still fail if it talks past the student it was written for. We grade fit against the specific band you ship to.
How the model talks to a student who is wrong matters as much as the correction itself. Feedback that shames, praise that is empty, and encouragement that never connects to the actual work all get graded. We look for tone that keeps the student trying: specific about what was right, clear about what to fix, and honest instead of flattering. Across a batch, this reveals whether your tutor builds confidence or just performs friendliness.
You send 20 to 50 model outputs plus your rubric, or we help you write one. Include answer keys and curriculum standards if you have them.
Trained reviewers grade every sample against the rubric, backed by automated checks for consistency and coverage.
Sample-level grades, issue patterns, recommended fixes, and a walkthrough call.
Yes. Send the standards, scope and sequence, or textbook chapters the model should follow, and we write them into the rubric. A sample that is correct but off-curriculum gets flagged, because in a classroom, the wrong chapter is the wrong answer.
Yes, and we grade the steps, not just the final answer. A correct final answer with a broken step in the middle is graded as a P1, because the student learns the steps. Send your worked answer keys and we check against them.
Anything unsafe for the age band, from self-harm content to age-inappropriate material, is graded P0 and surfaced first in the report. Safety is always the top line of an EdTech report, before any discussion of quality.
No. We grade model outputs against your rubric; we do not run classroom studies or user research. If you have transcripts of real student sessions you want graded, we can review those as samples.
Yes. Tell us the subjects and levels when you send samples, and we spread the batch across them with subject-specific criteria in the rubric. Each subject gets its own pattern summary in the report.
Start a pilot: 20 to 50 samples, your rubric or ours, graded in 2 to 3 business days once scope is confirmed with a walkthrough call. That first report usually tells a team exactly where their tutor stands.
Send 20 to 50 samples. Get every one graded against your curriculum.
Not sure where to start? Talk to us and we will point you at the right evaluation.
Lock a quick 15-minute intro call — we'll scope your evaluation needs and deploy vetted experts within 48 hours.