In Praise of Cannibalization: Why Your Best-Performing Treatment Should Get Less Traffic, Not More

Takeaway…

The accept rate you measure for a single treatment is not a property of that treatment. It is a property of the treatment and the customers your policy chose to show it to. The “winning” arm usually wins because a good contextual bandit quietly handed it a favourable audience. Roll it out to everyone and you are evaluating it on the customers it was never meant to see — and the number collapses.

Expecting the best-performing arm to get the most traffic is asking your contextual bandit to stop being contextual. If one arm really does deserve all the traffic, you didn’t need personalisation in the first place.


Last night, a substitute won the World Cup.

Ferran Torres came off the bench in the 62nd minute, and in the second half of extra time — against an Argentina side by then down to ten men and running on empty — he swept in the only goal of the final. His first of the whole tournament. On for less than an hour, decisive, a nation’s hero by midnight.

So here is the tempting conclusion: this is your best striker, start him every game. Ninety minutes, every match, from the first whistle.

Of course not. Absolutely not. The value Torres delivered was never a property of Torres. It was a property of Torres plus the situation his manager dropped him into: late, against tired legs and then ten men, with the game stretched wide open. Start him for ninety minutes against a fresh, organised back line and that match-winning ratio evaporates — not because he got worse, but because you changed the population of situations he was asked to perform in.

(Argentina, for their part, spent the evening proving a neighbouring point: if you decide the plan is to turn a football match into a wrestling bout, you have optimised for the wrong objective, and you lose the thing that actually mattered. But that is a measurement trap for another day.)

I keep meeting the super-sub’s twin in decisioning.

The most natural mistake in arbitration

Picture a team running NBA with several treatments live under a contextual policy. One treatment is, on paper, the clear winner — much the highest accept rate of the lot. Reasonable people look at that, and reasonably conclude: this is our best asset, let’s give it the most traffic.

Of course not. Absolutely not. And it’s worth being precise about why not, because the instinct is so natural that it’s almost never questioned.

That treatment very likely earned its accept rate by finding its way to a small but highly responsive sub-group. That is not an accident or a bug — it is the entire job of a contextual bandit. The policy routed the customers who love that treatment towards that treatment. The observed accept rate is therefore conditioned on a hand-picked audience. Serve it to everyone and you are now measuring it on the majority who were, correctly, routed elsewhere. The overall accepts don’t go up. They go down.

Two numbers that get conflated

Here is the confusion in one line. When people look at a per-arm accept rate, they think they are seeing this:

“How well would this treatment do if I gave it to everybody?”

But what they are actually seeing is this:

“How well did this treatment do on the specific customers my policy decided to send it?”

These are completely different quantities. The first is a property of the treatment. The second is a property of the treatment married to the allocation. A contextual policy is engineered to make the second number flattering — that’s success, not signal about the first.

A worked example, because numbers settle arguments

Two segments. S1 is 20% of the base and highly responsive to treatment A. S2 is the other 80%. Two treatments, A and B. Conditional accept probabilities:

Treatment A Treatment B
S1 (20%) 0.50 0.10
S2 (80%) 0.05 0.20

A competent contextual bandit learns to send A to S1 and B to S2. Now look at what you’d see on the treatment scorecard:

  • A, shown only to S1, has an observed accept rate of 50%.
  • B, shown only to S2, has an observed accept rate of 20%.

A is the runaway winner. So we roll A out to everyone:

  • A to the whole base = 0.20 × 0.50 + 0.80 × 0.05 = 14%.

The winner, given the traffic it “deserves,” delivers 14%. Meanwhile the policy we were about to override:

  • Contextual routing = 0.20 × 0.50 + 0.80 × 0.20 = 26%.

We would have almost halved our accepts by promoting the winner. And here’s the kicker for the provocateurs in the room: roll out the loser instead —

  • B to the whole base = 0.20 × 0.10 + 0.80 × 0.20 = 18%.

The treatment with the worst standalone number beats the treatment with the best standalone number when each is forced on the full population. If a single per-arm accept rate can invert the ranking like that, it has no business driving a reallocation decision.

cannibalization ~ contextual bandit

Regular readers will recognise the machinery, because it’s the same trap I wrote about in the action cannibalization paper — just viewed from the other side of the table.

In the cannibalization story, you launch a new action, total accepts rise, and yet the measured lift of an existing action falls. Everyone panics: the new action is cannibalising the old one. But nothing was stolen. The bandit simply started routing the high-propensity customers to whichever action suited them best, and the old action was left holding a less responsive residual audience. Its standalone number dropped precisely because the system got better at personalising.

That is the whole point:

cannibalization ~ contextual bandit

Here is the same idea from the action side, with the numbers small enough to check by hand. Before B exists, Action A serves a mixed audience of ten and converts 40% of them. Add B, and B draws away A’s most responsive customers — so A is now judged only on the less responsive customers left behind:

Nothing was stolen. Action A’s accept rate falls from 40% to 20% purely because its most responsive customers moved to a better-fitting action — while total accepts rise from 4 to 5. The falling number is the system getting better at personalising, not worse.

They are not two problems. They are one mechanism wearing two vocabularies. “Cannibalization” is the word we use when the falling number belongs to an action we already had. “The winning arm doesn’t scale” is the word we use when the flattering number belongs to an arm we’re tempted to promote. In both cases the per-item metric is confounded by which customers that item was allocated, and the confound is the bandit doing its job.

If you insist that every action defend its standalone lift, you are outlawing the very routing that creates value. You are asking the bandit to stop personalising so that your scorecard reads nicely. Don’t.

So are per-arm metrics useless?

No — dangerous, not useless, which is a distinction worth keeping. A per-arm accept rate is a perfectly good monitoring signal. It tells you an arm is alive, that it’s being served, that it isn’t quietly broken. What it cannot do is answer a counterfactual — “what happens if I move traffic here” — because it is silent about the population you’d be moving traffic from.

The metric that answers the reallocation question is the policy-level one: how much does the whole squad win, not how good is the ratio of the man you keep dropping into the perfect moment. Evaluate the policy, not the arm. If you must reason at arm level, reason about conditional performance per context, never the pooled marginal — which, incidentally, is exactly the pooled-versus-weighted AUC argument in different clothing.

And keep the diagnostic in your back pocket: if one arm genuinely deserves all the traffic — if forcing it on everyone really does beat the contextual policy — then congratulations, you’ve discovered that personalisation has no signal in this problem. That’s a real and useful finding. It’s just not the one people think they’re making when they promote the winner.


So: should you start the super-sub every week?

Only if the defence is always tired and always a man short — that is, only if the situation Torres thrived in is the situation everyone is in. In real customer bases it never is. The substitute’s ratio, the winning treatment’s accept rate, and the cannibalised action’s collapsing lift are all the same illusion: a number that looks like a property of one thing, but is really a property of a careful match between that thing and the moments you chose for it.

Measure the match. Not the man.


Wonder where the reference to spherical footballs came from? See Andy’s great article: Beyond PxV: Accounting for Non-Spherical Customers in Next-Best-Action

Otto, this is an excellent article, particularly because it challenges an assumption that feels intuitively correct until you stop and examine it.

In my experience, the single biggest factor separating successful NBA programmes from unsuccessful ones is whether an organisation genuinely escapes campaign thinking. Many teams implement sophisticated decisioning technology but continue to manage it as though it were a campaign management system. The result is a constant struggle between human control and machine optimisation.

The instinct that the best-performing treatment should receive more traffic comes directly from that campaign mindset. The Industry has spent decades teaching marketers to identify winners, allocate more budget, and scale them up. In a contextual decisioning environment, however, the apparent winner is often winning because the arbitration strategy has successfully learned where and when it is most effective. As you say, the outcome is a property of both the treatment and the context in which it was selected.

I see a similar challenge with traditional champion-challenger thinking. In classic testing, the loser has little value and is eventually culled. In NBA, however, today’s loser may be exactly the right treatment for a specific customer, in a specific situation, at a specific moment. A treatment that loses globally may still be a champion locally. Removing it can reduce optionality and ultimately weaken the system’s ability to personalise.

More broadly, this requires organisations to let go of a deeply ingrained belief that good decisions come from experienced marketers sitting around a table deciding priorities. I sometimes refer to this as the GOBSAT approach: Good Old Boys Sat Around a Table. The priorities may be informed by experience, politics, budgets, campaign calendars and occasional measurement, but they remain largely opinion-driven.

NBA demands a different mindset. Humans still play an essential role in defining objectives, value, constraints and guardrails, but the machine is better placed to determine which action should be presented to a particular customer at a particular moment. The goal is no longer to identify the single best treatment. The goal is to identify the best treatment for this customer right now.

That is why I particularly like your call to rethink cannibalisation. In traditional marketing it sounds like something to be feared. In NBA it is often evidence that the decisioning strategy is working as intended. If a treatment only performs well because the AI has learned exactly when and where to deploy it, forcing it onto everyone else isn’t exploiting a success, it is destroying the very context that made it successful in the first place. Excessive manual weighting results in auto-cannibalisation - autosarcophagy. (Sorry to those still eating breakfast).

So, great article. It captures one of the most important conceptual shifts organisations must make if they want to move from campaign management to true one-to-one decisioning.

Very nice article Otto and good lessons on NBA, football/soccer and wrestling.

I love your last sentence: measure the match, not the man. When you measure the match - goals, or in the example above, total accepts, you would notice that the overall KPI went up. The more detailed KPIs are still relevant, for instance to better understand drivers rather than just success, but should always be looked at in the context of overall performance.

The only preaching to-the-choir caveat I would like to make is that the above might not be the nr 1 problem facing teams. It is just a hunch, yet given an interaction many propositions are often weeded out by well-intended eligibility/suitability/applicability rules and contact policies, which essentially result in hard filtering/exclusion rules. If we would state ‘please define under what situations a customer would not be allowed to get the NBA recommendation even if potentially interested’ rules would get a lot looser.

It is partially also due to the fact that we don’t give people other options to provide input. In my view this would best not be defined as hard eligibility rules (unless it is a hard policy exclusion) but more as a way to provide a hypothesis. One way to operatonalize this to allow for rules / scorecards that become customer analytical inputs rather than exclusion filters. They could be persisted for a while, with an automated process weeding them out if they havent shown consistent value.

[PS after a chat with you it appears that contact policies are often the key drop-off point. There are some simple solutions, potentially get rid of contact policies on inbound, and make sure that you have a good set of Interaction History summaries as candidate predictors to measure the softer relationship between history and response]

But I sense some more great blogs coming from your end that might go into this topic. Tx for your posts!

Great article, I love how you clearly illustrate that contextual bandit is not a cannibal and nothing gets lost when introducing personalized actions and treatments, in fact much more is gained