§ MethodLast updated Aug 15, 2026 · ScatterAI

How we verify

ScatterAI builds and independently verifies production AI agents — systems that talk to your customers and cannot afford wrong answers.

How do you stop an AI agent from saying something wrong to a customer?

Every outbound reply is graded before it is delivered. A model-graded gate scores each draft against a written rubric; replies below threshold never ship — they route to a human instead. The customer only ever sees messages that passed. We call this the score-before-send gate.

How do you measure conversation quality without trusting the AI to grade itself?

Self-graded AI is how teams fool themselves, so we manufacture independence. Two isolated reviewers read the same conversations against the same written standard — injected verbatim, never paraphrased — and cannot see each other's output. Only findings both reviewers independently report survive to a human adjudicator. We call this the dual-blind audit.

What happens when you swap the model or edit a prompt?

It takes an exam first. Every human adjudication becomes a permanent test item, so the answer key grows with every review cycle. A new model, prompt, or rule set must match the incumbent's catch rate on confirmed errors before it ships. Report volume is never the acceptance metric; catch rate on known errors is. We call this the answer-key exam.

Where do your failure labels come from?

From the repairs, not the symptoms. We derive quality taxonomies backward from the levers a team can actually pull — a label exists only if it is wired to an executable fix, and the label's definition doubles as that fix's acceptance test. We call this the lever-backward taxonomy.

How do we know the method itself works?

We test it the way we test your AI. The method passes two acceptance rungs: a cold read by its author, then a blind execution by an independent agent on an invented scenario seeded with planted defects. Every revision ships with a dated changelog.

The numbers we publish — and where they come from

Every number on this site carries its denominator and its origin. Client measurements stay with clients; what we publish is measured on our own assets:

  • 3 / 3 planted defects caught — an independent agent blind-executed our audit method on an invented test scenario it had never seen, seeded with three known defects.
  • 27 documented method gaps found and patched across two adversarial acceptance rounds (9 by the author's own round, 18 by the independent round); every fix versioned in a dated changelog.
  • 24 adjudicated human verdicts in the answer-key exam set of our own research pipeline — the exam any grader change must re-pass before shipping.