We measured our own product until it failed.
It asked AI personas a question and aggregated their answers into a result. It generated 11,636 responses across 31 surveys and it worked exactly as designed. That turned out not to be the question that mattered.
4.8pp
Spread across reruns
7.7pp
Over-predicted yes
0.178
Brier, all questions
14.2pp
Mean absolute error
The instrument was steady.
Rerunning the same question produced nearly the same number every time — a standard deviation of 4.8pp, which is actually below the 11.1pp floor that binomial sampling alone would put on it. By any ordinary software standard the thing was working.
This is the trap. Stability feels like correctness, and a dashboard that returns the same answer twice looks trustworthy. It only means the ruler does not wobble. It says nothing about where zero is.
Spread across reruns
4.8pp
less than half the sampling floor
Distance from truth
7.7pp
band never reaches zero
Fig. 1 — Both bands are measured, on one scale. The reruns agree far more tightly than chance alone would give, and the whole band still misses zero.
It leaned yes, everywhere we looked.
Against 35 questions with known real-world answers, the personas said yes 7.7pp more often than people actually do. Two unrelated probes found the same lean without being designed to: a neutral-framing check at 6.4pp, and a framing-asymmetry check at 8.2pp. Three methods, one direction. That is a property of the instrument, not noise.
Fig. 2 — Each line is how far that answer missed by. Points below the diagonal said yes more often than people did.
And the edge was in the wrong half.
Overall it beat both naive baselines, which reads well until you split the corpus. On famous questions — the kind whose answer may sit in the training data — it was clearly better than the base rate. On preference and behaviour questions, the kind a customer would actually pay to ask, the base rate was essentially as good. We were selling the second kind.
All questions n=35
beats base rate by 0.067
Leakage-prone n=16
beats base rate by 0.060
Preference / behavior n=19
base rate is within 0.023
Brier score, lower is better. Shorter bar wins.
Fig. 3 — Split by question type, the advantage over the base rate is concentrated in exactly the questions whose answers could already be known.
What we did about it.
We took the word calibrated off the site, closed paid plans, and stopped new surveys. The whole battery cost 266 model calls and about $0.133 to run, which is roughly nothing against the months it would have taken to find this out from churn.
The product still runs. Existing results stay reachable and the code and data are kept. It is closed because we could not defend the number, not because it broke.
What is still open.
Everything above is retrospective, and grading your own homework after the fact is the weakest form of evidence there is. So before any outcome was known we locked 44 predictions behind a hash. 0 have resolved. The rest settle on their own schedule, and when they do we will not be able to move the answer.
sha256, locked 2026-06-22
ccdaac0b36e98af2d9ab7a91a88919e0c2358d34aadc69f87b7a5f4d8441887f