For years my position on multi-armed bandits in digital experimentation was short: don’t. A bandit assumes you pull a lever and find out what happened. On a website you pull the lever, and the customer wanders off, comes back on Thursday on a different device and buys the following Tuesday. I wrote that position up a while ago and I stood by it.

I’ve softened. Not because the problem went away, but because the tooling has started to deal with it properly, and because “you can’t use bandits” was hiding a more useful statement: you can, and here is exactly what you pay.

What you pay is the part almost nobody explains. The moment a bandit changes the traffic split, the arms stop containing the same mix of new and returning users, and a conversion rate read across them is biased. The standard fix is to compare only like-for-like cohorts. That fix is correct, and it comes with a bill of its own: every cohort has to finish converting before it tells you anything, so the faster you move traffic, the less you know when you move it.

This article is about that bill.

What a bandit assumes, and what a website does

A multi-armed bandit shifts traffic towards whichever variant is doing best while the test is still running. Thompson sampling, the algorithm both GrowthBook and Optimizely use for conversion metrics, sends each arm a share of traffic in proportion to the probability it is the best one. Less traffic wasted on losers, more conversions banked during the test. That is the whole appeal, and it is a real one.

The algorithm was designed for outcomes that arrive straight away. Show an ad, get a click or don’t. Each observation is complete the moment it is made, so the estimate for each arm is always built from finished data.

Ecommerce conversion is not like that. Somebody is assigned to a variant on Monday. They might buy on Monday. They might also come back on Wednesday, again on Saturday, and buy the following week. Until their conversion window closes, that user is an unfinished observation. They count in the denominator today and may or may not join the numerator later.

In a fixed 50/50 test this doesn’t matter. Both arms hold the same proportion of unfinished users at every moment, so whatever is missing is missing equally from both sides and the comparison stays fair. The trouble starts when the split moves.

An A/A test that reads 15% apart

Here is the smallest example I can build. It is arithmetic, not data from a live test, and every number in it can be checked with a calculator.

  • 1,000 new users a day for 14 days. Two arms, A and B, which are identical. Both truly convert 5% of users.
  • Of the users who will convert, 30% do it on the day they arrive and 10% on each of the next seven days. So a cohort is finished seven days after it was assigned.
  • Assignment is sticky. Once a user is in an arm, they stay there.
  • Days 1 to 7 run at 50/50. On day 8 the bandit, reacting to an early wobble, moves to 20/80 in favour of B and stays there.

Now read the pooled conversion rate for each arm at the end of day 14.

Arm AArm B
Users from week 1 (finished)3,5003,500
Users from week 2 (still converting)1,4005,600
Share of users still converting29%62%
Conversions recorded so far217343
Conversion rate as read4.43%3.77%
True conversion rate5.00%5.00%

Arm B reads 15% worse than arm A. Nothing is different about arm B. It has simply been handed four times as many users who haven’t finished converting yet.

Who is in each arm on day 14 Both arms truly convert at 5.00% Arm A · 4,900 users · 29% still converting reads 4.43% Week 1 · 3,500 · finished Week 2 · 1,400 Arm B · 9,100 users · 62% still converting reads 3.77% Week 1 · 3,500 · finished Week 2 · 5,600 · still converting Same product, same users, same true rate. The pooled read says arm B is 15% worse.

Users whose conversion window has closed. Users assigned in the last week, some of whom will convert but haven't yet. The arm the bandit favoured holds four times as many of them.

That is return-user bias in its plainest form. The arm that lost traffic is now dominated by older users, the ones who have had time to come back and buy. The arm that gained traffic is dominated by people who arrived yesterday. You are comparing a population of mostly returning visitors against a population of mostly new ones and calling the difference a treatment effect.

Three things that make it worse

The direction depends on how you count. Per user, as above, the up-weighted arm looks worse because its users are younger. Per session, the down-weighted arm looks better for a related reason: a bigger share of its sessions are return visits, and return visits almost always convert at a higher rate than first visits. Either way the read is confounded by who is in the arm, not what the arm does.

The bandit acts on its own distortion. A fixed test with a biased read gives you one wrong number at the end. A bandit feeds the number back into the next allocation. In the example, the arm it just promoted now looks like a loser, so it pulls traffic back, which shifts the age mix the other way. The weights move in response to the weights.

Switching off sticky assignment is not an escape. If returning users are re-rolled under the new weights, the age mix problem goes away and a worse one arrives: the same person sees both variants, and the purchase gets credited to whichever one they happened to land on last. GrowthBook’s documentation is direct about this: without sticky bucketing, “users could be re-assigned when they return to the experiment.”

There is a fourth problem underneath all of these, which is that sample means in adaptive experiments are biased even when outcomes are instant. Shin, Ramdas and Rinaldo showed in a 2019 paper that sampling more from arms that look good “induces a negative bias” in the estimates. That one is real but small next to the maturity effect above, and I mention it mainly so nobody thinks fixing delay fixes everything.

This is a close relative of Simpson’s paradox, and it is the reason I spent years saying bandits don’t belong on websites. The second tab covers the fix that changed my mind, and what the fix costs.


The point

I used to say you can’t use bandits where outcomes are delayed. The accurate version is narrower and more useful.

When a bandit moves the traffic split, the arms stop holding the same mix of new and returning users. Read a pooled conversion rate across them and you get a number driven by who is in each arm. In the worked example, two identical arms read 15% apart.

Comparing only within cohorts fixes that. GrowthBook does it by weighting across update periods. Optimizely’s Epoch Stats Engine does it by stratifying across periods of constant allocation. Both are right to.

But a cohort only tells the truth once it has finished converting. Update more often than your customers convert and the bandit is steering on early buyers, who are not a fair sample of all buyers. Wait for cohorts to finish and the bandit has little of the test left to steer. In a two-week test with a seven-day window and weekly updates, it has none.

So the question to ask before switching a bandit on is not which platform or which algorithm. It is how long your customers take to convert, compared with how long you plan to run the test. If the first number is small next to the second, a bandit will earn its keep. If it isn’t, the split you want is 50/50.


Whichever way that comes out, the assignment has to be sticky, consistent across visits and cheap to change, which is an infrastructure question before it is a statistics one. Conversion Workers handles it at the edge: the variant is decided before the page renders, stays with the user when they return through a sticky cookie, and the traffic split is configuration, not code. If you’d like to know whether a bandit could steer a useful share of your traffic at all, get in touch. Measuring your conversion lag is a small piece of work, and it settles the question either way.