How does the app know which buzz works, or which weight to set? It tests. It splits its users into two groups without telling them. One group gets the old version. The other gets a change. Then it counts minutes in each. People call this an A/B test: version A against version B.
Group A, minutes a day31 minutes
Group B, minutes a day33 minutes
Made-up result. B wins, so B ships.
Big apps run hundreds of these at once. You are in dozens right now. The colour of a button, the order of the tabs, the weight on shares, how many ads a session. Whatever wins more minutes, or more taps, or more money, goes to everyone. Whatever loses is quietly dropped.
Group A31 minutes a day
Group B33 minutes a day
Group A gets the old feed and spends 31 minutes a day. Group B gets a new rule and spends 33. What happens next?
The new rule goes to everyone.
Right. Two more minutes across millions of people is a lot of ad slots. The test asked one question, which version gets more minutes, and B answered it.
Nothing. Two minutes is too small to matter.
Not yet. Two minutes on 31 is about 6 in 100 more, and it repeats every day for every user. Small differences at scale are exactly what these tests are for.
Both stay, and users pick.
Not yet. Users are never asked. The point of the test was to avoid asking; people say one thing and do another, and the test counts what they did.
Which version gets more minutes?counted
vsvs
Which leaves people better off?no counter
Notice what the test asks and what it does not. It asks: which version gets more minutes? It does not ask: which version leaves people better off? The second question has no counter. So the test cannot answer it, and a thing that cannot be counted cannot win a test.
Move the control to set how many tests the app runs, one after another. The rule is made up: every test keeps a version that won two in a hundred more time.
Minutes a day at the start30 minutes
After the tests31 minutes
1 testAfter 1 test, each keeping the version that won two in a hundred more time, a 30-minute day is about 31 minutes.
Minutes a day at the start30 minutes
After the tests33 minutes
5 testsAfter 5 tests, each keeping the version that won two in a hundred more time, a 30-minute day is about 33 minutes.
Minutes a day at the start30 minutes
After the tests37 minutes
10 testsAfter 10 tests, each keeping the version that won two in a hundred more time, a 30-minute day is about 37 minutes.
Minutes a day at the start30 minutes
After the tests45 minutes
20 testsAfter 20 tests, each keeping the version that won two in a hundred more time, a 30-minute day is about 45 minutes.
Minutes a day at the start30 minutes
After the tests81 minutes
50 testsAfter 50 tests, each keeping the version that won two in a hundred more time, a 30-minute day is about 81 minutes.
No single test made a big difference. After fifty, the day has nearly tripled. Why?
Each test kept its small win, and the wins multiply.
Right. Two in a hundred, fifty times over, is not a hundred in a hundred. It is about a hundred and seventy, because each gain sits on top of the last. Nobody chose the tripling. The tests did, one small step each.
The fiftieth test was a big one.
Not yet. Every test in this rule adds the same two in a hundred. The size comes from adding on top of adding, not from any one test.
It would not. People have only so much time.
Not yet. True, and a real curve flattens. But it flattens at hours a day, not minutes, which is where fifty small wins got it.
Move the control.
In the testabout 700,000
vsvs
Told beforehandnone
Real example. In 2012 Facebook tested about 700,000 people without telling them. For a week, some saw fewer happy posts from friends and some fewer sad ones. Then it measured what they wrote. Those shown fewer happy posts wrote slightly sadder posts. The study came out in 2014, and many people were angry to learn of it.
Group A total30,000 minutes in a day
Group B total33,000 minutes in a day
An app has 2,000 users in a test. Half get the new version. Group A averages 30 minutes a day, group B averages 33. Over one day, how many more minutes did group B spend in total?
minutes
A thousand people, three minutes each.
Right. Three thousand minutes, or fifty hours, from one small change in one day in a group of a thousand. Now send the winner to a billion users.
A company wants to test each of these. Which has a counter the app can count, and which does not?
One card at a time. Tap the pile it belongs to.
Card 1 of 6
Does a red badge bring more people back?
Does the feed make people kinder?
Do people watch longer with the sound on by default?
Do people feel lonelier after an hour?
Does a bigger share button get more shares?
Do people learn more from the app than they used to?
Does a red badge bring more people back? → Has a counter
Not yet. Has a counter. Opens are logged, so the app can compare opens with the badge and without it.
Does the feed make people kinder? → No counter
Not yet. No counter. Nothing on a phone logs kindness, so there is no number for the two groups to be compared on.
Do people watch longer with the sound on by default? → Has a counter
Not yet. Has a counter. Minutes watched are the app's favourite number, and the sound setting is one switch to flip.
Do people feel lonelier after an hour? → No counter
Not yet. No counter. Loneliness leaves no trace in the logs; an hour spent lonely and an hour spent happy look the same.
Does a bigger share button get more shares? → Has a counter
Not yet. Has a counter. Every share is a logged tap, and a button's size is a one-line change to test.
Do people learn more from the app than they used to? → No counter
Not yet. No counter. Learning happens inside a person, and the app cannot count what went in.
Loggedgets tested
vs
Inside a personnever asked
Right. A test needs a number the phone can log. Opens, minutes, taps, shares: logged. Kindness, loneliness, learning: not logged. So the first list gets tested and improved, and the second list is never asked.
Minutes before100 minutes, before = 100
Minutes after120 minutes, before = 100
A homework app adds a leaderboard and a daily streak. Two months later its minutes are up 20 in 100. Its makers say the app is better. What can the test actually tell them?
Nothing. Tests with streaks are not real tests.
Not yet. It is a real test of the thing it counted. The limit is not the method but the counter, which was time and not learning.
That the minutes went up. Whether anyone learned more was never counted.
Right. The test's counter was minutes. Minutes rose, so the test succeeded on its own terms. Learning was not measured, so nothing the test says can be about learning.
That students learned more, since they used it more.
Not yet. More minutes is more minutes. A leaderboard can raise time on an app without raising what is learned, and the test cannot tell those apart.
Lesson complete
Every button, colour and feed rule was tried on some users first, and the version that won more minutes stayed.