Skip to main content

A one-person studio
that ships small apps.

One app at a time, each doing a single thing, 50 of them live on a Korean super-app platform. The figure shows what they have added up to, counted the two ways the platform counts it. Unevenly: 3 of the 50 produced about 73% of the first opens and 69% of the active days, and the median app produced 93.

Portfolio150 days to 2026-09-02
through2026-09-02
total7 daysfirst opens20,806861return days6,295209active days27,1011,070
010k20k30k01k2k3krunning total27.1k20.8k103per 7 days, trailingreleases50 apps liveAprMayJunJulAugSep
Complete calendar months, portfolio totals. First opens and return days are both summed across the 50 apps and neither is a count of people.
MonthFirst opensReturn daysActive days
2026-056,1411,6647,805
2026-065,1321,6076,739
2026-071,7941,1442,938
2026-082,0938352,928

Fig. 1. Neither count is people: both are summed across the 50 apps, and active days counts somebody again on every day they come back. At least 69.7% of first opens were never followed by a return day, and the 1.3 active days per open belongs to the largest apps: the median app is 1.09. Across complete months first opens fell from 6,141 in May to 2,093 in Aug, and return days from 1,664 to 835. Both halves fell, acquisition faster. What moved either is not in this data.

Selected work

What we have shipped.

Audience research

Retired

A survey tool that asked AI personas a question and aggregated the answers.

Responses generated
11,636
Surveys run
31
Retrospective Brier
0.18
ICC, noise-adjusted
~1.0

Consistent, but not accurate enough to sell. We closed it before anyone had paid for it.

2026Read the case study

Mini-apps

Live

Small single-purpose tools built for a Korean super-app platform, shipped as a portfolio rather than as one bet.

Apps live
50
Active days, summed
27,101
First opens, summed
20,806

Built one at a time on a platform that already had the audience, which is the opposite of how the product above was launched. Both figures are summed across the apps and neither is a count of people: a first open is somebody counted once per app they ever open, an active day is that same person counted again on every day they come back. The 6,295 days between the two are the whole of the return behaviour, 1.3 active days per first open, which is a weighted mean rather than the typical app: the median app is 1.09. Across complete months first opens fell from 6,141 in May to 2,093 in Aug, and return days from 1,664 to 835: both halves fell, acquisition faster. At least 69.7% of first opens were never followed by a return day. It is concentrated too: 3 of the 50 produced about 73% of the opens and the median app produced 93. They run inside that platform's app, so there is no public link to give, and the individual apps stay unlisted.

2026 onward

All work, what shipped and what happened to it.

How we measure

Consistent is not the same as correct.

Anyone can rerun a model until the number stops moving. That only proves the instrument is steady. We check the harder thing: whether the steady number is the right one.

Error266 calls
sampling floor±11.1+7.7truth-50+5+10+15+20percentage points from truth

Spread across reruns

4.8pp

less than half the sampling floor

Distance from truth

7.7pp

band never reaches zero

Fig. 2. Both bands are measured, on one scale. The reruns agree far more tightly than chance alone would give, and the whole band still misses zero.

Repeat runs of the same question had a standard deviation of just 4.8pp, under half the 11.1pp that sampling noise alone would produce. Steadier, in fact, than independent sampling permits, which means the personas inside one run are not independent voices.

It was also wrong in a fixed direction. Across 35 questions it over-predicted yes by 7.7pp, and two unrelated probes reproduced the same lean at 6.4pp and 8.2pp.

Worse, the answer moved with the packaging. Simply reordering the options shifted it by 12.5pp on average.

See the full result

Reworded question

5.4pp shift

max 17.3pp

Flipped framing

8.2pp shift

Reordered options

12.5pp shift

max 25.9pp

Sealed

locked 2026-06-22

Everything above is retrospective, and grading your own homework after the fact is the weakest evidence there is. So before any outcome was known we wrote down 44 predictions, hashed them, and published the hash. 7 have resolved since. 13 more future events and 24 human surveys have not. Because the hash came first, the answers could not be moved to suit them.

First result

7 resolved, July 2026

All 7 were results from one football tournament, on outcomes that had not been determined when the predictions were sealed.

52.9%

the average chance the panel gave them

14.3%

the share that happened, one of 7

Scored, it lands close to a coin either way: 0.284 setting aside the 35.7% of personas who answered neither yes nor no, 0.174 counting them as no, and at 7 questions the interval around both contains the coin’s 0.250, even computed as though the seven were independent, which one tournament is not. Part of the gap was fixed before any match was played, because four of the seven asked whether a given team would win and they cannot all be right.

Every question, both conventions, and why this is not a verdict yet

The check that needs no outcomes

Coherence4 exclusive outcomes
coherent max193.4 total050100150200percent, summed
  • France52.3%
  • Spain50.8%
  • England44.4%
  • Argentina45.9%

Fig. 3. Four sealed questions asked whether a given team would win the same tournament, and only one team can. A coherent set of probabilities has to fit under 100. Sampling noise does not reach this: the standard error on the sum is 12.6 points and the overshoot is 7.4 of them.

That part does not depend on who won, on how many questions have resolved, or on which convention you score under: four teams that cannot all win were given 193.4% between them, or 121.6% counting the maybe answers as no. The predictions contradict each other before reality is consulted.

Working on something small?

We are a studio of one, so we take on very little. If the problem is narrow and you want to know whether it actually worked, write to us.