Beyond AUC: Why Success Rate Is the Metric That Actually Measures NBA Success

Formula 1 teams invest heavily in wind tunnel testing. The aerodynamic numbers matter because they correlate with race-day performance. But if a team topped every wind tunnel benchmark and still finished tenth on Sunday, nobody would call that season a success.

Customer Decision Hub has a similar dynamic.
AUC is the wind tunnel metric.
Success rate is race-day performance.

AUC is valuable, and you should monitor it. But in live Next-Best-Action, the question is not whether a model looks good in validation. The question is whether customers actually say yes when decisions are made in real channels, under real constraints, at real scale.

That is what success rate measures.

Takeaway

In Customer Decision Hub, AUC tells you whether a model can separate likely responders from non-responders. Success rate tells you whether your Next-Best-Action program is actually creating outcomes in production. If you care about business impact, success rate is the primary metric.

What AUC is good for, and what it is not

AUC captures discrimination quality: how well a model ranks positive outcomes above negative ones. In CDH terms, how well propensity aligns with actual behavior.

That makes AUC useful for:

  1. Model development and challenger testing.
  2. Comparing candidate predictors on historical data.
  3. Monitoring model learning quality when production outcomes are not yet attributable.

But AUC is often over-interpreted. It is a geometric statistic, the area under a ROC curve. Even though it is expressed numerically between 0.5 and 1, it is not a probability of business success.

So when someone says, “Our AUC is great, we’re winning,” the correct response is: maybe. Show the production success rate.

Why success rate is the KPI for live NBA

In production, the objective is straightforward: maximize meaningful positive outcomes. Clicks, accepts, conversions, completions.

Success rate measures exactly that:

Success Rate = Number of successful outcomes / Number of decisions shown

This makes it operationally superior as a primary business metric because it is:

  1. Directly tied to customer behavior.
  2. Directly tied to value realization.
  3. Comparable across policy and strategy changes.
  4. Sensitive to whether personalization is working where it matters, at decision time.

A model can have a strong AUC and still underperform if it is not improving the decisions that are actually being surfaced.

Shadow mode explains the distinction clearly

CDH already reflects this split in Prediction Studio monitoring:

  1. In ‘shadow’ mode, a model learns but does not drive decisions. Success rate is not applicable as a causal production metric for that model.
  2. In active mode, the model influences who gets what. Now success rate becomes the proof point.

So the hierarchy is practical:

  1. Before activation: rely on AUC and related diagnostics.
  2. After activation: optimize and judge primarily by success rate.

Precision and recall: useful, but context-dependent

A common question is why not use precision and recall as metrics of the model’s accuracy.

The answer is scope. Precision and recall are ranking/retrieval metrics. They matter when there is a set of alternatives to rank and retrieve, such as multi-offer surfaces where several offers are shown.

In many Next-Best-Action implementations, only the single best action is shown. In that setup, the problem is binary response prediction:

  • Will the customer accept this offer or not?

In that case, success rate is the natural business metric. There is no meaningful retrieval set to evaluate, so precision and recall are simply not applicable.

Precision and recall become valid only when you intentionally show a ranked set of offers, for example a top-3.

When multiple offers are in play:

  1. Precision asks: of the offers shown, how many were relevant?
  2. Recall asks: of all relevant offers available, how many were surfaced?

Interesting connection: in many practical settings, Precision@1 aligns with success rate at the top position. That is another way of saying the business outcome at rank 1 is still the essential scoreboard.

NBA is not just response prediction, it is action differentiation

In Next-Best-Action, success depends on choosing the best action, not merely predicting generic click tendency.

This matters because some predictors inflate response propensity in a broad sense but do little to distinguish among candidate actions. For example, “customer clicked before” may be strongly predictive of future clicking overall, yet weak for deciding which action should win arbitration right now.

So even with a healthy AUC, you can fail to maximize NBA effectiveness if your model is not strong at action-level differentiation. Success rate at the decision layer reveals whether that differentiation is truly working.

Control groups turn success rate into a rigorous performance test

CDH control design provides one of the cleanest reality checks available in production:

  1. Test group uses propensity-driven personalization.
  2. Control group receives an eligible random action (still respecting policy and eligibility).

Now compare outcomes:

  1. Control success rate shows baseline performance without propensity-based targeting.
  2. Personalized success rate shows performance with model-driven relevance.
  3. The delta is lift in success rate.

Lift = Success Rate Personalized − Success Rate Control

In the monitoring of CDH this is divided by Success Rate Control, to express a relative Lift.

This is the strongest practical evidence that your model is not just statistically elegant, but economically useful. It is also a direct quantitative measure of usefulness, and therefore an explicit expression of business value.

Final thought

AUC is an excellent laboratory instrument.
Success rate is the business truth.

Lift in success rate goes one step further: it shows how much the model improves outcomes over a baseline, so it captures the business value created by personalization.

Formula 1 offers the same lesson. A team can post impressive simulator numbers, but if race-day results are below expectation, as Dutch fans have felt at times with Max Verstappen this year :wink: , the season is judged by points and podiums, not by test-bench charts.

In CDH Next-Best-Action, the model exists to improve real decisions. The program that consistently improves success rate, especially against a valid control, is the program that is truly learning, truly personalizing, and truly delivering value.


Two cautions from recent discussions

  1. Pooled AUC can be inflated and hide action-level variation. A weighted per-action view is how AUC should be measured. https://forums.pega.com/t/your-85-auc-is-probably-a-lie-heres-the-math/12926

  2. A single action’s success rate may decrease while overall portfolio success rises. That is not necessarily failure. It can be optimal reallocation of traffic toward stronger total-system outcomes.
    https://forums.pega.com/t/in-praise-of-cannibalization-why-your-best-performing-treatment-should-get-less-traffic-not-more/12813

Great post @Ivar_Siccama, thanks for this! Really glad to see this conversation around success rates getting more attention. You’re absolutely right that success rates or downstream conversion behaviors are what actually matters to the business.

I’d emphasize the lift of success rates, as you hinted at already: while higher success rates are obviously better, the tricky part is defining “good” without context. Is 1% good? 5%? It totally depends on the channel, organization, use case. Lift is more grounded —how much better your approach performs against a baseline (the previous implementation, random selection, etc.). That gives you something concrete to measure against.

That said, the bar matters. A 5% lift over baseline? Probably not moving the needle enough. But a few hundred percent over random prioritization? That’s more like it—though of course you need enough optionality in your offer pool to make that realistic. (Optionality is its own topic, and worth a future post, stay tuned…)

Nice post!

It is indeed easy to get lost in the weeds of KPIs, and it is important to start from the top down. In that sense, framing this as separating outcomes and drivers can be useful. Both need to be measured, but it is important to start with the outcome,

Overall success/accept rate is a good example of an outcome, and model quality such as AUC as a driver (but certainly not the only one). What a predictive model can do is to identify customers where the likelihood to accept a specific NBA x is say 1.5-3 times higher than for the average/a randomly selected customer (a lift of 1.5-3 in the top scoring 10%).

But if say the accept rate for x for a random customer is say 1% and for the other NBA y it is 5%, the predictive power of the model becomes irrelevant. The general appeal of a NBA is a more important driver in this instance.

And even models that would just output the overall accept rate for each NBA would still be very useful even though the models wouldn’t be predictive / have an AUC of 0.50 (!).

In other words the goal of the models is to rank NBAs on the likelihood of a positive outcome, and well-tuned models can have this additional lift of 1.5-3 over the likelihood for a random customer for a specific NBA.

It was also the original thinking behind the core ‘starry night’ bubble chart in Prediction Studio:

  • NBAs with a high success rate and AUC are the real stars.
  • NBAs with low AUC and high success rate are the ‘don’t care segment’ from a data science perspective as the outcome is already great (which also means it gets harder to obtain lift).
  • Likewise NBAs with a poor outcome / success rate but a high AUC are not the modelers problem: it is simply not an appealing proposition (you could try loosening constraints drastically)

For a data scientist everything is interesting to investigate, especially the machine learning stuff. But from a business perspective, outcome KPIs should guide where to focus.