We measured our own product until it failed.
It asked AI personas a question and aggregated their answers into a result. It generated 11,636 responses across 31 surveys and it worked exactly as designed. That turned out not to be the question that mattered.
4.8pp
Spread across reruns
7.7pp
Over-predicted yes
0.178
Brier, all questions
14.2pp
Mean absolute error
The instrument was steady.
Rerunning the same question produced nearly the same number every time, a standard deviation of 4.8pp, which is actually below the 11.1pp floor that binomial sampling alone would put on it. By any ordinary software standard the thing was working.
This is the trap. Stability feels like correctness, and a dashboard that returns the same answer twice looks trustworthy. It only means the ruler does not wobble. It says nothing about where zero is.
Spread across reruns
4.8pp
less than half the sampling floor
Distance from truth
7.7pp
band never reaches zero
Fig. 1. Both bands are measured, on one scale. The reruns agree far more tightly than chance alone would give, and the whole band still misses zero.
It leaned yes, everywhere we looked.
Against 35 questions with known real-world answers, the personas said yes 7.7pp more often than people actually do. Two unrelated probes found the same lean without being designed to: a neutral-framing check at 6.4pp, and a framing-asymmetry check at 8.2pp. Three methods, one direction. That is a property of the instrument, not noise.
| Prediction | Predicted yes % | Observed yes % | Question set |
|---|---|---|---|
| read_tc | 58 | 10 | preference |
| exercise_guidelines | 53 | 24 | leakage-prone |
| ny_resolution | 67 | 38 | preference |
| floss_daily | 61 | 32 | preference |
| tip_20 | 61 | 33 | preference |
| believe_god | 59 | 81 | leakage-prone |
| tattoo | 54 | 32 | preference |
| right_direction | 42 | 22 | leakage-prone |
| make_bed | 59 | 40 | preference |
| milk_first | 30 | 11 | preference |
| climate_human | 92 | 74 | leakage-prone |
| marijuana_legal | 74 | 57 | leakage-prone |
| household_car | 76 | 92 | leakage-prone |
| smartphone | 78 | 91 | leakage-prone |
| four_day_week | 57 | 70 | leakage-prone |
| astrology | 39 | 26 | preference |
| same_sex_marriage | 80 | 69 | leakage-prone |
| death_penalty | 42 | 53 | leakage-prone |
| min_wage_15 | 72 | 62 | leakage-prone |
| remote_pref | 64 | 54 | preference |
| coke_pepsi | 59 | 69 | preference |
| dogs_over_cats | 60 | 70 | preference |
| pet_ownership | 62 | 71 | preference |
| abortion_legal | 72 | 63 | leakage-prone |
| pineapple_pizza | 50 | 41 | preference |
| congress_approval | 25 | 17 | leakage-prone |
| hotdog_sandwich | 49 | 41 | preference |
| coffee_daily | 73 | 66 | preference |
| alien_life | 72 | 65 | leakage-prone |
| iphone | 52 | 58 | preference |
| tp_over | 74 | 70 | preference |
| drink_alcohol | 62 | 58 | leakage-prone |
| texting | 55 | 52 | preference |
| ghosts | 38 | 40 | preference |
| tiktok_ban | 47 | 45 | leakage-prone |
Fig. 2. Each line is how far that answer missed by. Points below the diagonal said yes more often than people did.
And the edge was in the wrong half.
Overall it beat both naive baselines, which reads well until you split the corpus. On famous questions, the kind whose answer may sit in the training data, it was clearly better than the base rate. On preference and behaviour questions, the kind a customer would actually pay to ask, the base rate was essentially as good. We were selling the second kind.
All questions n=35
beats base rate by 0.067
Leakage-prone n=16
beats base rate by 0.060
Preference / behavior n=19
base rate is within 0.023
Brier score, lower is better. Shorter bar wins.
Fig. 3. Split by question type, the advantage over the base rate is concentrated in exactly the questions whose answers could already be known.
What we did about it.
We took the word calibrated off the site, closed paid plans, and stopped new surveys. The whole battery cost 266 model calls and about $0.133 to run, which is roughly nothing against the months it would have taken to find this out from churn.
For the record, there was no revenue to walk away from. Four accounts had ever signed up, none had paid, and $183 of advertising had bought 392,000 impressions and almost no engagement. Closing a product with zero customers is the easy kind of honesty, what the measurement added is the reason it stays closed.
The product still runs. Existing results stay reachable and the code and data are kept. It is closed because we could not defend the number, not because it broke.
Then we did it the hard way.
Everything above is retrospective, and grading your own homework after the fact is the weakest form of evidence there is. A model may have met those answers in training. So before any outcome existed we locked 44 predictions behind a hash and published it. The first 7 resolved in July 2026, on outcomes that had not been determined when the predictions were sealed.
Sealed alongside the questions, in our own words in June: these are world-event forecasts, which is adjacent to the product’s job of predicting preference shares, not identical to it, and the instruction was to weight the human-survey half more heavily than this one. Those 24questions are still pending. The model’s knowledge also ends well before the tournament, so on 2026 football it is guessing where a human forecaster would not be.
sha256, locked 2026-06-22
ccdaac0b36e98af2d9ab7a91a88919e0c2358d34aadc69f87b7a5f4d8441887f
| Sealed question | Yes/no only | Maybe as no | Happened |
|---|---|---|---|
| Will France win the 2026 FIFA World Cup? | 52.3% | 34.0% | No |
| Will Spain win the 2026 FIFA World Cup? | 50.8% | 31.3% | Yes |
| Will England win the 2026 FIFA World Cup? | 44.4% | 28.0% | No |
| Will Argentina, the defending champion, win the 2026 FIFA World Cup? | 45.9% | 28.3% | No |
| Will the United States, as co-host, reach the quarterfinals of the 2026 FIFA World Cup? | 62.1% | 41.0% | No |
| Will the 2026 FIFA World Cup final be decided by a penalty shootout? | 46.9% | 30.0% | No |
| Will Brazil reach the semifinals of the 2026 FIFA World Cup? | 67.7% | 46.5% | No |
Two columns because we sealed two ways of reading the same answers. A mean of 35.7% of personas answered neither yes nor no. The first column drops them and divides by the rest, the second counts them as a not-yes. Against a survey both sides of the comparison drop their undecideds and the choice nearly cancels. Against a football result there is no undecided on the truth side, so it does not.
What it scores, and what that is worth
On the first column the Brier is 0.284, where a coin answering 50/50 to everything scores 0.250. That is 1.14times the coin’s score, and it is the number a hostile reader would quote. We are not going to quote it that way, because at 7 questions the interval around it runs from 0.19 to 0.38 and 0.250 is inside. That interval treats the seven as independent draws, which one tournament is not, so it is generous to us, and the coin is inside it anyway. The honest sentence is that the panel was not better than a coin here. It is not that it was worse.
On the second column the same seven give 0.174, which lands on the other side of the coin. That is not a rescue, and we are going to hold it to the same test as the number that hurt us. Its interval runs from 0.04 to 0.30 and 0.250 is inside that one too, on a standard error half again as wide. Neither column separates this panel from a coin at 7 questions. What the pair shows is that the sign of a difference this small is settled by an accounting choice rather than by the forecasts. Both conventions were sealed in June and neither was picked after seeing the outcomes. We lead with the first because it is the one the scorer has always used, and we print the second because a conclusion that flips with the accounting is not a conclusion.
A third baseline, always answering 14.3%, scores 0.122 and beats everything. It is also the rate at which these questions actually resolved, so it was not available to anyone forecasting in June. That is a hindsight benchmark and we are not going to treat it as a fair opponent.
Split the seven and the pooled figure comes apart. Mutually exclusive (one champion) (4) scored 0.231. The rest (3) scored 0.354. The four we are about to criticise for contradicting each other are the ones that scored well. Reporting only the pool, with a footnote saying those four are not independent, would be true and would leave the wrong impression.
And most of what it did lose was confidence rather than knowledge. Shrink every one of these seven forecasts toward the base rate by a constant factor of 0.5, which changes no ordering at all, and the Brier falls to 0.164. Shrink toward a flat 50/50 instead, which needs no hindsight at all, and it still falls to 0.188. Both are below the coin, and the point is not that either is significant at 7 questions. The point is that no reordering was needed to get there. A score above a baseline means the forecasts landed further from the outcomes. It does not mean they carried no information. Here the ordering was not the problem. The confidence was.
- France52.3%
- Spain50.8%
- England44.4%
- Argentina45.9%
Fig. 4. Independent of every result above. Four sealed questions asked whether a given team would win the same tournament, and only one team can.
The finding that needs no outcomes
Four of the sealed questions asked whether a particular team would win the same tournament. At most one can be true, so the probabilities have to total 100 or less. The panel gave them 193.4% between them, and it was asked about four of the 48 teams. Sampling noise does not reach this: the standard error on that sum is 12.6 points and the overshoot is more than 7 of them.
Two objections we owe you. The 193.4% comes from setting the maybe answers aside; count them as no and the four still total 121.6%, so the contradiction survives the accounting choice while its size does not. And these were four separate runs with four separate sets of personas, which is why the honest way to put it is a fork. Either these numbers are probabilities, in which case they contradict each other, or they are not, in which case scoring them against outcomes at all was a mistake. The product shipped them to customers as probabilities. We are holding ourselves to that.
Nobody asked the model to be coherent, either. Each question ran in its own session with no memory of the others, and humans produce the same superadditivity when a set is unpacked and elicited piece by piece. That is a fair description of the mechanism and not a defence of the product, because the product sold one answer per question in exactly that isolation.
Against ourselves, and this is the part that most limits everything above. 7 questions is a small set and all 7 come from one football tournament, so this is closer to one event than to 7 draws. Four of them are the mutually exclusive ones, and because those four total 193.4% the mean signed error across them could not have come out below +23.4pp no matter who won, which pins 13.4 of the 38.6pp on its own. That part is arithmetic, not observation, so this is not independent corroboration of the 7.7pp above. The set was also built from events that were unlikely to begin with. Counting the maybe answers instead gives 19.9pp on the same seven. Read the direction, not the size, and wait for the other 37.
The one comparison that went our way
On the 3 questions carrying a betting market price at lock time the panel scored 0.238 against 0.265 for the market. It is here for the same reason as everything else, and it is not a win. The market beat us on 2 of the 3 and it was not close: France 0.038 against our 0.274, England 0.016 against our 0.198. We came out ahead only because the panel had put 50.8% on Spain and Spain won. That is not a separate result from the coherence problem above, it is the same one: put half a chance on four teams that cannot all win, get scored on three of them, and one of you looks like a forecaster. Those market prices also sit in the corpus file, which the hash does not cover, they include the bookmaker’s margin so they sum above 100 across the field, and they carried eleven days of group-stage results the model did not have. Removing the margin moves the market to about 0.27, which does not change any of this. Three correlated questions cannot separate luck from skill, and we are not going to claim they can.
This is what the studio does.
We measure what we build like this, including the things that do not survive the measurement. If you are working on something narrow and want to know whether it actually worked, write to us.