How we verify
ScatterAI builds and independently verifies production AI agents — systems that talk to your customers and cannot afford wrong answers.
How do you stop an AI agent from saying something wrong to a customer?
Every outbound reply is graded before it is delivered. A model-graded gate scores each draft against a written rubric; replies below threshold never ship — they route to a human instead. The customer only ever sees messages that passed. We call this the score-before-send gate.
How do you measure conversation quality without trusting the AI to grade itself?
Self-graded AI is how teams fool themselves, so we manufacture independence. Two isolated reviewers read the same conversations against the same written standard — injected verbatim, never paraphrased — and cannot see each other's output. Only findings both reviewers independently report survive to a human adjudicator. We call this the dual-blind audit.
What happens when you swap the model or edit a prompt?
It takes an exam first. Every human adjudication becomes a permanent test item, so the answer key grows with every review cycle. A new model, prompt, or rule set must match the incumbent's catch rate on confirmed errors before it ships. Report volume is never the acceptance metric; catch rate on known errors is. We call this the answer-key exam.
Where do your failure labels come from?
From the repairs, not the symptoms. We derive quality taxonomies backward from the levers a team can actually pull — a label exists only if it is wired to an executable fix, and the label's definition doubles as that fix's acceptance test. We call this the lever-backward taxonomy.
How do we know the method itself works?
We test it the way we test your AI. The method passes two acceptance rungs: a cold read by its author, then a blind execution by an independent agent on an invented scenario seeded with planted defects. Every revision ships with a dated changelog.
The numbers we publish — and where they come from
Every number on this site carries its denominator and its origin. Client measurements stay with clients; what we publish is measured on our own assets:
- 3 / 3 planted defects caught — an independent agent blind-executed our audit method on an invented test scenario it had never seen, seeded with three known defects.
- 27 documented method gaps found and patched across two adversarial acceptance rounds (9 by the author's own round, 18 by the independent round); every fix versioned in a dated changelog.
- 24 adjudicated human verdicts in the answer-key exam set of our own research pipeline — the exam any grader change must re-pass before shipping.