Simpson's paradox
One hospital is better for every kind of patient, and still has worse results overall. Choose a hospital, look inside the numbers, then move the patients around and watch the verdict flip. About five minutes.
1. Choose a hospital
You need an operation, and there are two hospitals nearby. Each treated 1,000 patients for it last year. Here is how many survived.
Which hospital would you choose?
2. Look inside the numbers
Not every patient arrives in the same state. Split each hospital's patients into mild and severe cases.
| Patients | Riverside | Hilltop |
|---|---|---|
| Mild | 776 of 800 (97%) | 198 of 200 (99%) |
| Severe | 120 of 200 (60%) | 520 of 800 (65%) |
| All | 896 of 1,000 (90%) | 718 of 1,000 (72%) |
Hilltop is better for mild patients, and better for severe patients. It is better for every patient you could be, and yet its overall number is far worse.
Look at who each hospital treats. Most of Riverside's patients are mild. Most of Hilltop's are severe, perhaps because it's the one ambulances go to. The overall figure mixes two very different groups in very different amounts.
3. Move the patients
Hilltop's survival rates within each group stay fixed: 99% for mild cases and 65% for severe. What if it treated a different mix of patients?
If Hilltop treated the same mix as Riverside (1 in 5 severe), which would look better overall?
With Riverside's mix, Hilltop's overall survival would be about 92%, better than Riverside's 90%. Nothing about Hilltop's care changed. Only the mix of patients did.
Hilltop looks worse overall once more than about 28% of its patients are severe cases. Its real share is 80%.
4. Why: an average of averages
An overall rate is an average of the group rates, weighted by how many patients are in each group. Change the weights and the average moves, even when every group rate stays the same.
Riverside = 0.8 × 97% + 0.2 × 60% = 89.6%
Hilltop = 0.2 × 99% + 0.8 × 65% = 71.8%
The paradox needs two things at once. Something (here, severity) affects the outcome, and the groups being compared contain very different amounts of it. Then the comparison that matters is within each group, like with like.
Splitting isn't always the right move, though. Severity was decided before anyone chose a hospital, so it's fair to hold it fixed. If the thing you split by is itself caused by the treatment, such as blood pressure measured after taking a blood-pressure drug, splitting hides the very effect you're trying to measure. Split by what came first.
5. Something new
A university has two departments. Overall, 53% of the men who applied were admitted, and 37% of the women.
| Department | Men admitted | Women admitted |
|---|---|---|
| Engineering | 480 of 800 (60%) | 130 of 200 (65%) |
| English | 50 of 200 (25%) | 240 of 800 (30%) |
| All | 530 of 1,000 (53%) | 370 of 1,000 (37%) |
What explains the overall gap?
Each department admitted women at a higher rate than men. Most women applied to English, which turns most people away. Most men applied to Engineering, which admits most of them. The overall gap comes from where people applied, not from how each department treated them.
This pattern turned up in real graduate admissions data at the University of California, Berkeley, in 1973. It's why "compare within groups" is the first thing to try when an overall number surprises you.