For years my position on multi-armed bandits in digital experimentation was short: don’t. A bandit assumes you pull a lever and find out what happened. On a website you pull the lever, and the customer wanders off, comes back on Thursday on a different device and buys the following Tuesday. I wrote that position up a while ago and I stood by it.
I’ve softened. Not because the problem went away, but because the tooling has started to deal with it properly, and because “you can’t use bandits” was hiding a more useful statement: you can, and here is exactly what you pay.
What you pay is the part almost nobody explains. The moment a bandit changes the traffic split, the arms stop containing the same mix of new and returning users, and a conversion rate read across them is biased. The standard fix is to compare only like-for-like cohorts. That fix is correct, and it comes with a bill of its own: every cohort has to finish converting before it tells you anything, so the faster you move traffic, the less you know when you move it.
This article is about that bill.
What a bandit assumes, and what a website does
A multi-armed bandit shifts traffic towards whichever variant is doing best while the test is still running. Thompson sampling, the algorithm both GrowthBook and Optimizely use for conversion metrics, sends each arm a share of traffic in proportion to the probability it is the best one. Less traffic wasted on losers, more conversions banked during the test. That is the whole appeal, and it is a real one.
The algorithm was designed for outcomes that arrive straight away. Show an ad, get a click or don’t. Each observation is complete the moment it is made, so the estimate for each arm is always built from finished data.
Ecommerce conversion is not like that. Somebody is assigned to a variant on Monday. They might buy on Monday. They might also come back on Wednesday, again on Saturday, and buy the following week. Until their conversion window closes, that user is an unfinished observation. They count in the denominator today and may or may not join the numerator later.
In a fixed 50/50 test this doesn’t matter. Both arms hold the same proportion of unfinished users at every moment, so whatever is missing is missing equally from both sides and the comparison stays fair. The trouble starts when the split moves.
An A/A test that reads 15% apart
Here is the smallest example I can build. It is arithmetic, not data from a live test, and every number in it can be checked with a calculator.
- 1,000 new users a day for 14 days. Two arms, A and B, which are identical. Both truly convert 5% of users.
- Of the users who will convert, 30% do it on the day they arrive and 10% on each of the next seven days. So a cohort is finished seven days after it was assigned.
- Assignment is sticky. Once a user is in an arm, they stay there.
- Days 1 to 7 run at 50/50. On day 8 the bandit, reacting to an early wobble, moves to 20/80 in favour of B and stays there.
Now read the pooled conversion rate for each arm at the end of day 14.
| Arm A | Arm B | |
|---|---|---|
| Users from week 1 (finished) | 3,500 | 3,500 |
| Users from week 2 (still converting) | 1,400 | 5,600 |
| Share of users still converting | 29% | 62% |
| Conversions recorded so far | 217 | 343 |
| Conversion rate as read | 4.43% | 3.77% |
| True conversion rate | 5.00% | 5.00% |
Arm B reads 15% worse than arm A. Nothing is different about arm B. It has simply been handed four times as many users who haven’t finished converting yet.
Users whose conversion window has closed. Users assigned in the last week, some of whom will convert but haven't yet. The arm the bandit favoured holds four times as many of them.
That is return-user bias in its plainest form. The arm that lost traffic is now dominated by older users, the ones who have had time to come back and buy. The arm that gained traffic is dominated by people who arrived yesterday. You are comparing a population of mostly returning visitors against a population of mostly new ones and calling the difference a treatment effect.
Three things that make it worse
The direction depends on how you count. Per user, as above, the up-weighted arm looks worse because its users are younger. Per session, the down-weighted arm looks better for a related reason: a bigger share of its sessions are return visits, and return visits almost always convert at a higher rate than first visits. Either way the read is confounded by who is in the arm, not what the arm does.
The bandit acts on its own distortion. A fixed test with a biased read gives you one wrong number at the end. A bandit feeds the number back into the next allocation. In the example, the arm it just promoted now looks like a loser, so it pulls traffic back, which shifts the age mix the other way. The weights move in response to the weights.
Switching off sticky assignment is not an escape. If returning users are re-rolled under the new weights, the age mix problem goes away and a worse one arrives: the same person sees both variants, and the purchase gets credited to whichever one they happened to land on last. GrowthBook’s documentation is direct about this: without sticky bucketing, “users could be re-assigned when they return to the experiment.”
There is a fourth problem underneath all of these, which is that sample means in adaptive experiments are biased even when outcomes are instant. Shin, Ramdas and Rinaldo showed in a 2019 paper that sampling more from arms that look good “induces a negative bias” in the estimates. That one is real but small next to the maturity effect above, and I mention it mainly so nobody thinks fixing delay fixes everything.
This is a close relative of Simpson’s paradox, and it is the reason I spent years saying bandits don’t belong on websites. The second tab covers the fix that changed my mind, and what the fix costs.
Compare like with like
The bias described in the first tab comes from pooling users who were assigned under different traffic splits. The fix is to stop pooling them.
Split the test into periods during which the allocation was constant. Within one period, both arms were filled at the same time, under the same split, so their users are the same age and have had the same opportunity to return. Compare the arms inside each period, then combine the per-period comparisons into one estimate. Every time the bandit changes the weights, a new period starts.
In the worked example, that means comparing week-1 A against week-1 B, and week-2 A against week-2 B. Week 1: 5.00% against 5.00%. Week 2: both arms have recorded exactly the same fraction of their eventual conversions, so they read the same as each other. The phantom 15% disappears.
This is stratification, and it is what both major platforms do in some form. Optimizely’s documentation puts the principle well: “bias due to Simpson’s paradox cannot occur over periods of constant traffic allocation.”
I think this is the right approach. It is also where most explanations stop, and that is a pity, because the interesting part comes next.
A cohort that hasn’t finished is measuring something else
Within a period, both arms are the same age. That makes the comparison fair. It does not make it complete.
If a cohort is two days old, comparing its arms tells you which variant converts more people within two days. That is a legitimate metric. It is not the one you care about, and the two can disagree.
Suppose variant B adds urgency messaging. It doesn’t persuade anyone who wasn’t going to buy. It just gets them to buy sooner. Final conversion is 5% in both arms, but in B half of the eventual buyers convert on day one instead of 30%.
| Age of cohort | Arm A reads | Arm B reads | Apparent lift |
|---|---|---|---|
| Day of assignment | 1.50% | 2.50% | +67% |
| 3 days old | 3.00% | 3.57% | +19% |
| 7 days old (finished) | 5.00% | 5.00% | 0% |
A bandit reading that cohort on day one sees a 67% winner and moves traffic accordingly. The comparison was perfectly fair. The variant is worth nothing. The bandit has optimised for speed of purchase and nobody asked it to.
So a cohort is only decision-grade once its conversion window has closed. Before that, any up-weighting is a bet that early converters are representative of all converters, which is exactly the assumption that return-user behaviour breaks.
The trade-off: cadence against maturity
Here is the part I think people don’t appreciate. You now have two dials, and they pull against each other.
How often the bandit updates. Each update closes one cohort and opens another. Frequent updates mean the bandit reacts quickly, and they also mean many small cohorts.
How long a cohort takes to mature. This isn’t a dial at all. Your customers set it. If people take a week to buy, a cohort takes a week to finish.
Day the cohort is assigned. Maturing: conversions still arriving. Finished: safe to compare.
Run a two-week test with a seven-day conversion window and look at what is actually known on the last day. Cohorts from days 1 to 7 are finished. Cohorts from days 8 to 14 are still converting. Those are the same cohorts the bandit was steering. The traffic it moved is precisely the traffic you can’t read yet, and the honest readout is on day 21, a week after the test stopped.
That leaves three ways to run it, and all three cost something.
Update often, on everything. The bandit moves daily using whatever data exists, finished or not. It is responsive and it is steering on early converters, as in the urgency example. This is what most people picture when they say “bandit”.
Update often, on finished cohorts only. Now every decision rests on complete data, and the data is a week old. The first cohort is collected on day 1 and finishes on day 8, so the first informed reallocation happens at the start of day 9, and it rests on one day of traffic. In the example that is 1,000 users and about 25 conversions per arm, well under the 40 conversions per variation GrowthBook recommends before a bandit starts moving weights.
Update slowly, on finished cohorts only. Weekly cohorts are big enough to trust. The first one is collected by day 7 and finishes on day 14. In a two-week test, that is the day the test ends. The bandit never reallocates anything. You have run an A/B test with extra steps.
Share = 1 − (conversion window + update cadence) ÷ test length. The first cohort has to be collected, then finish converting, before any reallocation rests on complete data.
The chart is the whole argument in one picture. The share of traffic a bandit can steer on finished evidence is roughly one minus (conversion window plus update cadence) divided by test length. If the window is a large fraction of the test, there is nothing left to steer. If the window is close to zero, almost everything is.
Two more items on the bill
Unequal splits are expensive to read. The precision of a comparison inside a cohort depends on how evenly the cohort was split. A cohort split 90/10 carries 36% of the information of the same cohort split 50/50, because the small arm is the bottleneck. A bandit that is working well produces lopsided cohorts by design, so the periods where it was most confident are the periods that contribute least to the final estimate.
The cohorts are not the same kind of people. A Monday cohort and a Saturday cohort behave differently. Stratifying handles that for the comparison between arms, but only on the assumption that the difference between variants is steady over time. Optimizely’s documentation says so openly: time variation that affects the arms differently “remains an open area of research.”
None of this makes the cohort approach wrong. It makes it a trade. You have swapped a biased estimate for an unbiased one that arrives late, costs precision, and limits how much of the test the bandit can actually use. How GrowthBook and Optimizely each make that trade is the subject of the next tab.
Two platforms, two philosophies
I read the bandit documentation for both GrowthBook and Optimizely for this piece. They agree on the algorithm and disagree on almost everything about what a bandit is for. Everything in this tab comes from their published docs as of October 2026, not from inside knowledge of either product.
| GrowthBook | Optimizely | |
|---|---|---|
| Algorithm | Thompson sampling | Thompson sampling for binary metrics, epsilon-greedy for numeric |
| Update cadence | You choose. Minimum 15 minutes, recommended daily or longer | ”The MAB model is updated hourly” |
| Before the first update | Exploration window: at least 100 users per variation required, 40 conversions per variation recommended | Not described in the docs I read |
| Floor on a losing arm | Every variation keeps at least 1% of traffic | Not described in the docs I read |
| Returning users | Sticky bucketing “generally” advised | A user profile service is recommended “to ensure sticky bucketing” |
| Changing-split bias | Results “weighted over bandit update periods” | Epoch Stats Engine stratifies by “periods of constant allocation”, built for Stats Accelerator |
| Inference on a bandit | Shown, with a warning that “these results may be biased” | None: “MABs ignore statistical significance” |
GrowthBook: a bandit you can still read
GrowthBook treats a bandit as an experiment with a moving split. It tries to keep the estimates usable, and it tells you where they aren’t.
On the bias from changing weights, the docs say GrowthBook is “weighting across periods to deal with the fact that users entering your experiment on different days have different behavior”, and then immediately adds “but it is still a limitation of the Bandit approach.” That is the cohort fix from the second tab: one period per update, compared within, combined across.
On delay, the guidance is the cadence trade-off stated in two sentences. The decision metric should have “a short conversion window”, and “ideally your users convert in a window much shorter than your bandit’s update cadence.” On update frequency: “we recommend setting your update cadence to daily or even longer”, because longer windows “reduce the likelihood that a fluky day of traffic will cause undesirable weight updates.”
And on when not to bother, the docs say that with a “long sales cycle, or if you care about long-term effects, a standard Experiment may be better”, and that bandits can do worse than a standard experiment “when there are only two arms.”
What I couldn’t find documented is a maturity gate: a setting that holds a cohort out of the weight calculation until its conversion window has closed. The docs handle delay by telling you to pick a fast metric and a slow cadence. That is sound advice, and it is advice rather than a guardrail. If you point a daily bandit at a seven-day purchase metric, nothing I read suggests the tool will stop you.
Optimizely: a bandit is not a test
Optimizely’s position is blunter, and I have some respect for it. A bandit there is an optimiser. It “only focuses on the variations’ empirical mean”, there is no statistical significance on the results page, and the docs say that is deliberate: it “avoids confusion about the purpose and meaning of MAB optimizations.”
In other words, Optimizely doesn’t solve the readout problem for bandits. It declines to give you a readout. If you want inference with adaptive allocation, that is a different feature, Stats Accelerator, and that is where the Epoch Stats Engine lives. Epochs are periods of constant allocation, each weighted “proportional to the total number of visitors within that epoch”. Same idea as GrowthBook’s period weighting, with the maths published in more detail.
The part that worries me is the hourly update. An hourly cadence is fine when the outcome arrives within the hour: a click, a signup, an add-to-basket. For a purchase that takes days, an hourly bandit is the “update often, on everything” option from the second tab at its most extreme. Every decision is made on whoever converts fastest. The epoch documentation doesn’t discuss conversions that land after the epoch they were assigned in has closed, and at one-hour epochs that is most of them.
My earlier article gave Optimizely credit for sticky assignment and for only applying new weights to new users. That still stands. What I’d add now is that sticky assignment is what creates the age-mix problem in the first place. It is necessary, and it is not sufficient.
Which is better
For a metric with any meaningful delay, GrowthBook’s approach is the better one, for three reasons.
- You control the cadence. The single most important setting, relative to your conversion window, is yours to set. With a fixed hourly update it isn’t.
- It keeps the estimate and labels it honestly. A weighted read with a bias warning is more useful than no read, provided people take the warning seriously.
- The documentation tells you when to walk away. Long sales cycles, two arms, long-term effects. Most vendors don’t write that page.
Optimizely’s approach is the better one if you accept its framing: you have a short-lived campaign, an outcome that lands in minutes, and no intention of learning anything from the result. For a weekend promotion with four headlines, that is a perfectly good tool and an honest one.
Neither platform, as documented, does the thing I would most want: refuse to move traffic on a cohort that hasn’t finished converting. Both leave that to you. The last tab is about how to hold that line yourself.
When a bandit earns its place
Having spent three tabs on the costs, here is the other side. There are conditions where a bandit is the right tool, and they are easy to state.
A bandit works when the outcome arrives much faster than the bandit moves, and the bandit moves many times before the test ends. Put as a rule of thumb from the cadence chart in the second tab: conversion window plus update cadence should be a small fraction of the test length. If it’s under a fifth, a bandit can steer most of the traffic on finished evidence. If it’s over half, run an A/B test.
That gives a short list of good fits:
- Fast outcomes. Click-through on a banner, a headline, a recommendation slot, an email subject line. The outcome lands in the same session, so cohorts mature almost immediately and the maturity problem from the first tab barely exists.
- Many arms. Six headlines, not two layouts. With two arms a 50/50 split is already close to optimal and a bandit mostly adds bias.
- Short-lived decisions. A promotion that ends on Monday. There is no “ship the winner” afterwards, so the value is all in what you earn during the test, and the quality of the final read matters less.
- Cheap to be wrong about the size. You need to pick the best arm, not report a lift to finance.
And a short list of poor ones: considered purchases, anything B2B, subscription retention, revenue per user, and any test whose purpose is to learn something you will reuse.
How to run one on a slow metric anyway
Sometimes the metric is slow and the case for adapting is still good, usually because there are many arms and some of them are clearly bad. These are the rules I would hold to.
- Measure your conversion lag before you configure anything. Take last quarter’s converters and plot days from first visit to purchase. The number you need is the day by which 90% of eventual converters have converted. Most teams have never looked, and it is usually longer than they’d guess.
- Set the cadence no shorter than that lag, or steer on a faster proxy. Add-to-basket rate matures in a session. If you use a proxy, check on historical tests that it agrees with the real metric often enough to trust.
- Hold out unfinished cohorts. Only cohorts whose window has closed feed the weights. If your platform can’t do this, a long cadence is the nearest substitute.
- Keep assignment sticky and analyse per user. Per-session reads bring return-visit mix straight back in.
- Keep a floor under every arm. Ten per cent, not one. It costs a little reward and buys back a great deal of precision in the final read.
- Read the result a full window after the last user was assigned. Not on the day the test stops.
- Run an A/A bandit first. Two identical arms through the full setup. If it reports a winner, you have found your bias for the price of a fortnight.
Rules 2 and 3 will, on many sites, leave the bandit with very little room to adapt. That is the finding, not a failure of the method. It means the test should have been a fixed split, perhaps with sequential stopping rules so you can end it early when the evidence is clear. That gets most of the benefit people want from a bandit with none of the reassignment bias.
What I still don’t know
I have not run a controlled comparison of GrowthBook’s period weighting against a maturity-gated design on live traffic, so I can’t put a number on how much the difference matters in practice. The worked examples here are arithmetic. They show the mechanism and its direction. The size on your site depends on your conversion lag and how hard the bandit swings, and the only way to find that out is the A/A run in rule 7.
How Conversion Works can help
This is usually a three-step piece of work, and the first step is often the last.
Find out whether you have the problem. We measure your conversion lag, look at how your current tests allocate and read traffic, and tell you whether a bandit could steer a useful share of it. For most ecommerce sites on a purchase metric the answer is that it can’t, and a well-run fixed test is the better buy. That answer is worth having before anyone pays for an auto-allocation feature.
If it can, set it up so the read survives. Cadence matched to lag, maturity gating, a proxy metric validated against the real one, sticky assignment at the edge, and an A/A run to prove the setup is clean before a real test goes through it.
Then check it paid. Compare what the bandit earned against what a fixed split would have earned over the same period. If the difference is inside the noise, go back to fixed splits and keep the simpler system.
Who this isn’t for
- Teams testing clicks and headlines. Your outcomes are immediate. Turn the bandit on and enjoy it.
- Two-variant tests. Run an A/B test. Both platforms’ own docs agree.
- Sites that struggle to power a test at all. Splitting thin traffic into cohorts makes it thinner. The low-traffic playbook is a better use of the afternoon.
The point
I used to say you can’t use bandits where outcomes are delayed. The accurate version is narrower and more useful.
When a bandit moves the traffic split, the arms stop holding the same mix of new and returning users. Read a pooled conversion rate across them and you get a number driven by who is in each arm. In the worked example, two identical arms read 15% apart.
Comparing only within cohorts fixes that. GrowthBook does it by weighting across update periods. Optimizely’s Epoch Stats Engine does it by stratifying across periods of constant allocation. Both are right to.
But a cohort only tells the truth once it has finished converting. Update more often than your customers convert and the bandit is steering on early buyers, who are not a fair sample of all buyers. Wait for cohorts to finish and the bandit has little of the test left to steer. In a two-week test with a seven-day window and weekly updates, it has none.
So the question to ask before switching a bandit on is not which platform or which algorithm. It is how long your customers take to convert, compared with how long you plan to run the test. If the first number is small next to the second, a bandit will earn its keep. If it isn’t, the split you want is 50/50.
Whichever way that comes out, the assignment has to be sticky, consistent across visits and cheap to change, which is an infrastructure question before it is a statistics one. Conversion Workers handles it at the edge: the variant is decided before the page renders, stays with the user when they return through a sticky cookie, and the traffic split is configuration, not code. If you’d like to know whether a bandit could steer a useful share of your traffic at all, get in touch. Measuring your conversion lag is a small piece of work, and it settles the question either way.