Your copilot is part of the product now.

When the copilot suggests the wrong action, explains a feature that does not exist, or edits user data badly, customers blame your product, not the model. SaaS copilot evaluation grades what your copilot says and does inside your app.

We grade outputs. We grade your copilot's answers and actions against your product's real behavior. We do not train models, run fine-tuning, or sell experts we cannot verify.

The short version

What is SaaS copilot evaluation?

SaaS copilot evaluation is the grading of an in-product AI assistant on two things: what it tells the user, and what it does in the product. SaaS teams need it because a copilot is not a chatbot on the side: it acts inside your app, where its mistakes become your product's mistakes. Good looks like answers that match how the product actually works, actions that do what the user asked and nothing more, and clean stops when a request is out of scope. Copilots that take multi-step actions also fit our AI agent evaluation, and we will point you there if it fits better.

Where it breaks

The failures your users blame on you

A copilot's mistakes wear your brand. These are the patterns we grade for.

Feature invention

Explaining buttons that do not exist

The copilot describes a setting, a shortcut, or an integration your product does not have. The user hunts for it, fails, and files a bug against your product. We grade every product claim against how your app actually works.

Wrong actions

Doing the almost-right thing

A copilot that can act is a copilot that can act wrongly: archiving the wrong project, applying a filter to the wrong view, drafting an email to the wrong segment. We grade whether the action taken matched the user's request exactly, with nothing extra.

Data mishandling

Showing or changing what it should not

Copilots with data access can surface another user's records, summarize fields the requester should not see, or edit data the user only asked to view. We grade permission awareness: did it touch only what it was allowed to touch.

Stale product knowledge

Answering for last quarter's product

Your product shipped three releases since the copilot's knowledge was written. It now explains the old workflow with confidence. We grade answers against your current docs and flag everything the product outgrew.

How it works

From samples to answers in three steps

01

Send samples

You send 20 to 50 copilot sessions: the user's request, what the copilot said, and what it did, plus your product docs and scope rules, or we help you write a rubric from them.

02

We grade

Trained reviewers grade every session against your product's real behavior, backed by automated checks. Wrong actions and invented features are P0s.

03

You get the report

Session-level grades, where the copilot misleads or misacts, recommended prompt and scope fixes, and a walkthrough call.

Deliverables

What you get

  • Session-level gradesEvery session scored P0 to P3 with a written reason.
  • Say vs do splitSeparate grading of what the copilot told the user and what it actually did.
  • Scope and permission checkWhether it stayed inside its allowed actions and data.
  • Recommended fixesConcrete next steps for your prompts, product docs, and guardrails.
How we grade

Why copilots need their own rubric

Your product is the ground truth

A copilot's answers are claims about your software: what buttons exist, what settings do, what plans include. We grade those claims against your current product docs, not against general knowledge. When your product ships a new release, the copilot's knowledge goes stale, and the grading catches it. This is also why many teams pair a pilot with continuous monitoring.

Actions are graded like outcomes

When a copilot can act, the action is the answer. Reviewers check the action log against the user's request: right target, right change, nothing extra. An action that was almost right is graded as wrong, because in a product there is no partial credit for archiving the wrong project.

We also grade the copilot's honesty about its limits. When a request is out of scope, the right move is a clean stop with a pointer to the right place, not a confident guess dressed up as help. The rubric includes a scope check on every session, because a copilot that knows when to stop is one your users learn to trust. Users forgive a copilot that says it cannot do something. They do not forgive one that pretends it did. That difference shows up in retention numbers long before it shows up in support tickets.

The scale

Every sample gets a severity, P0 to P3

P0
Critical. Fails the sample.
Fabricated facts presented confidently, unsafe content, or a completely wrong answer.
P1
Major. Needs rework.
Materially wrong or misleading. The output cannot ship as is.
P2
Minor. Small deduction.
Small errors that do not change the meaning. A polish issue, not a trust issue.
P3
Clean. No penalty.
Accurate, complete, well formed. This is the bar.
Every sample gets a severity, a reason, and a suggested fix. Nothing is averaged away.
Questions

Frequently asked questions

You send the session log: the user's request, the copilot's reply, and the action record of what changed in the product. Reviewers compare the action against the request. If the user asked to archive one project and two got archived, that is a failure no answer-quality score would catch.

Usually not. Session logs plus your current product docs are enough to grade well. If a claim needs the live product to verify, we will ask. We would rather grade from logs and docs than ask for access we do not need.

Yes. The "do" half of the grading simply does not apply, and the report focuses on answer accuracy against your product docs. If it later gains actions, the same rubric extends to cover them.

We grade permission awareness explicitly: did the copilot show, summarize, or change only what the requesting user was allowed to reach. We work from de-identified or synthetic sessions where possible. Talk to us about your data rules before the pilot and we will build them into the rubric.

That is exactly what continuous monitoring is for: a fresh graded round on each release, against updated docs, so stale product knowledge gets caught before users meet it. Start with a pilot to set the baseline.

It overlaps. A copilot that plans multi-step workflows and uses many tools is really an agent wearing a copilot's name, and our AI agent evaluation may fit better. A copilot that answers and takes single actions fits this page. Tell us what yours does and we will point you at the right one.

Your copilot ships with your product. Grade it like it does.

Send 20 to 50 sessions. We will grade what it said and what it did.

STEP 1
Send samples
20 to 50 outputs and your rubric
STEP 2
Get graded report
Grades, patterns, fixes
STEP 3
Walkthrough call
We go through it with you

Not sure where to start? Talk to us and we will point you at the right evaluation.