Research Methods · Simplified
How Much Do Two Groups Overlap?
An effect size is a statement about two distributions sitting on top of each other. It is not a statement about whether there is an effect.
Before you look
Make a guess first
Two groups differ by what psychologists call a large effect: Cohen's d of 0.8. Out of a hundred people in the lower group, how many score above the average of the higher group?
The two distributions
The same gap, described four ways
Key terms
- Cohen's d
- The gap between the two means, measured in standard deviations. It has no units, which is what lets it be compared across studies that measured different things.
- Overlap
- The share of the two distributions that sits in the region they have in common. The shaded area on the figure.
- Probability of superiority
- Pick one person at random from each group. This is the chance that the one from the higher group scores higher.
- U3
- The percentage of the higher group that scores above the average of the lower group.
The solid curve is the higher-scoring group and the dashed curve the lower-scoring one. The shaded region is the part they have in common. There is no score axis here on purpose: none of these four numbers depends on the units.
What this shows
Large effects still overlap enormously
A d of 0.8 is conventionally called large. It leaves about 69 per cent of the two distributions in common, and picking one person from each group gets the ordering people expect only about 71 times in a hundred.
Key idea: Cohen's d is the gap between two means measured in standard deviations, exactly as a z-score measures one person against one distribution. Because it has no units it can be compared across studies, and because it is a ratio it is unchanged when both the gap and the spread are scaled together: two studies with nothing in common can produce the identical picture. All four numbers here describe that one picture. None of them says whether the effect is real, and none of them gets larger when you test more people. That is the whole reason effect sizes are reported alongside p-values rather than instead of them.
Three cautions. Cohen's labels of small, medium and large were offered as a rough fallback for researchers with nothing better, and he said so; they are not benchmarks, and what counts as a consequential effect depends entirely on the outcome and the cost of acting on it. The figures here assume both groups are normal with the same spread, which is what makes the four indices agree; where the spreads differ, d depends on which standard deviation you divide by and the overlap no longer follows from d alone. And a group difference licenses nothing about any individual: at d of 0.8 you would still be wrong nearly three times in ten guessing which of two people scored higher.
The longer version adds the sample-size control that separates an effect size from a test statistic, a matching challenge across four indices, and worked presets from published effects. It is at Cohen's d and Distributional Overlap in the main collection.