Winning every comparison and losing overall

Algorithm A beats algorithm B on easy images. It also beats B on difficult images. Surely A must have the better overall accuracy? That conclusion can fail when the algorithms are tested on different mixtures of cases. Simpson’s paradox shows how a correct calculation can support a misleading comparison.

Consider a deliberately constructed example. On easy images, A classifies nine out of ten correctly, giving 90% accuracy. B classifies eighty out of one hundred correctly, giving 80%. On difficult images, A gets 30% correct. B gets 20%.

A wins in both categories. Across all images, however, A has 39 correct answers out of 110, approximately 35.5%. B has 82 out of 110, approximately 74.5%. The overall ranking reverses, even though every percentage has been calculated correctly.

The explanation is the mixture of cases. Most of A’s images were difficult, while most of B’s were easy. An overall percentage is a weighted average of the category percentages. Here, each algorithm receives different weights, so the totals combine differences in performance with differences in the assignments.

A meaningful comparison needs a shared target population. Suppose the intended application contains equal numbers of easy and difficult images. Using the illustrative category rates, A’s expected accuracy is 60 per cent and B’s is 50 per cent. Both calculations now give the categories equal weight. These remain estimates. For otherwise comparable independent samples, using fewer observations increases sampling uncertainty.

The reversal itself is arithmetic. Deciding which comparison answers the real question takes further reasoning. Pearl’s discussion of Simpson’s paradox explains why the process generating the data matters, especially when the aim is to identify a causal effect. [1]

Dividing every dataset into smaller groups is therefore not an automatic solution. A grouping variable may itself be affected by the process under investigation. Conditioning on it can change the question or introduce bias. Useful subgroup analysis needs a reason for treating those groups as comparable.

For a performance table, the denominator deserves as much attention as the headline percentage. How many cases were assessed, how difficult were they, and do they represent the eventual workload? In this example, B’s impressive overall score reflects a much easier test. Once both systems face the same mixture, A’s advantage becomes visible. The apparent contradiction disappears when the weights are made explicit.

This is why apparent data should always be taken with a grain of salt until you can verify its truth through mathematics.

[1] Judea Pearl – Understanding Simpson’s paradox

Leave a comment