A one-person studio
that ships small apps.
One app at a time, each doing a single thing, 50 of them live on a Korean super-app platform. The figure shows what they have added up to, counted the two ways the platform counts it. Unevenly: 3 of the 50 produced about 73% of the first opens and 69% of the active days, and the median app produced 93.
| Month | First opens | Return days | Active days |
|---|---|---|---|
| 2026-05 | 6,141 | 1,664 | 7,805 |
| 2026-06 | 5,132 | 1,607 | 6,739 |
| 2026-07 | 1,794 | 1,144 | 2,938 |
| 2026-08 | 2,093 | 835 | 2,928 |
Fig. 1. Neither count is people: both are summed across the 50 apps, and active days counts somebody again on every day they come back. At least 69.7% of first opens were never followed by a return day, and the 1.3 active days per open belongs to the largest apps: the median app is 1.09. Across complete months first opens fell from 6,141 in May to 2,093 in Aug, and return days from 1,664 to 835. Both halves fell, acquisition faster. What moved either is not in this data.
Selected work
What we have shipped.
Audience research
RetiredA survey tool that asked AI personas a question and aggregated the answers.
- Responses generated
- 11,636
- Surveys run
- 31
- Retrospective Brier
- 0.18
- ICC, noise-adjusted
- ~1.0
Consistent, but not accurate enough to sell. We closed it before anyone had paid for it.
Mini-apps
LiveSmall single-purpose tools built for a Korean super-app platform, shipped as a portfolio rather than as one bet.
- Apps live
- 50
- Active days, summed
- 27,101
- First opens, summed
- 20,806
Built one at a time on a platform that already had the audience, which is the opposite of how the product above was launched. Both figures are summed across the apps and neither is a count of people: a first open is somebody counted once per app they ever open, an active day is that same person counted again on every day they come back. The 6,295 days between the two are the whole of the return behaviour, 1.3 active days per first open, which is a weighted mean rather than the typical app: the median app is 1.09. Across complete months first opens fell from 6,141 in May to 2,093 in Aug, and return days from 1,664 to 835: both halves fell, acquisition faster. At least 69.7% of first opens were never followed by a return day. It is concentrated too: 3 of the 50 produced about 73% of the opens and the median app produced 93. They run inside that platform's app, so there is no public link to give, and the individual apps stay unlisted.
All work, what shipped and what happened to it.
How we measure
Consistent is not the same as correct.
Anyone can rerun a model until the number stops moving. That only proves the instrument is steady. We check the harder thing: whether the steady number is the right one.
Spread across reruns
4.8pp
less than half the sampling floor
Distance from truth
7.7pp
band never reaches zero
Fig. 2. Both bands are measured, on one scale. The reruns agree far more tightly than chance alone would give, and the whole band still misses zero.
Repeat runs of the same question had a standard deviation of just 4.8pp, under half the 11.1pp that sampling noise alone would produce. Steadier, in fact, than independent sampling permits, which means the personas inside one run are not independent voices.
It was also wrong in a fixed direction. Across 35 questions it over-predicted yes by 7.7pp, and two unrelated probes reproduced the same lean at 6.4pp and 8.2pp.
Worse, the answer moved with the packaging. Simply reordering the options shifted it by 12.5pp on average.
See the full resultReworded question
max 17.3pp
Flipped framing
Reordered options
max 25.9pp
Sealed
locked 2026-06-22Everything above is retrospective, and grading your own homework after the fact is the weakest evidence there is. So before any outcome was known we wrote down 44 predictions, hashed them, and published the hash. 7 have resolved since. 13 more future events and 24 human surveys have not. Because the hash came first, the answers could not be moved to suit them.
First result
7 resolved, July 2026All 7 were results from one football tournament, on outcomes that had not been determined when the predictions were sealed.
52.9%
the average chance the panel gave them
14.3%
the share that happened, one of 7
Scored, it lands close to a coin either way: 0.284 setting aside the 35.7% of personas who answered neither yes nor no, 0.174 counting them as no, and at 7 questions the interval around both contains the coin’s 0.250, even computed as though the seven were independent, which one tournament is not. Part of the gap was fixed before any match was played, because four of the seven asked whether a given team would win and they cannot all be right.
Every question, both conventions, and why this is not a verdict yetThe check that needs no outcomes
- France52.3%
- Spain50.8%
- England44.4%
- Argentina45.9%
Fig. 3. Four sealed questions asked whether a given team would win the same tournament, and only one team can. A coherent set of probabilities has to fit under 100. Sampling noise does not reach this: the standard error on the sum is 12.6 points and the overshoot is 7.4 of them.
That part does not depend on who won, on how many questions have resolved, or on which convention you score under: four teams that cannot all win were given 193.4% between them, or 121.6% counting the maybe answers as no. The predictions contradict each other before reality is consulted.
Working on something small?
We are a studio of one, so we take on very little. If the problem is narrow and you want to know whether it actually worked, write to us.