Research Methods
Cohen's d and Distributional Overlap
An effect size is a statement about two whole distributions, not about two averages. The picture that matters is how much of them sit on top of one another — and it is almost always more than people expect.
Simulated — illustrative distributions, no real group data
Learning objective
By the end you should be able to read Cohen's d as overlap between two distributions, and say why sample size cannot move it.
About 25 minutes. Nothing you do here is saved or sent anywhere.
Groups are not people
Every index on this page describes two distributions. None of them says anything about any individual. Even at d = 1.0 — twice Cohen's "medium" — nearly a third of the two distributions still overlaps, and roughly one person in four drawn from the higher group scores below a person drawn from the lower one. Sentences of the form "group X are more Y" almost always overstate what an effect size supports.
- Estimate the overlap before you see it.
- Move the means apart and watch every index move together.
- Change one group's spread and watch d change without the means moving.
- Change the sample size and watch d refuse to move at all.
First, an estimate
Two fictional groups are compared on a 100-point task. Both have a standard deviation of 10; their means are 50 and 58. That is d = 0.80, which Cohen's rough labels would call a large effect.
The explorer
Two normal distributions on one axis. The shaded common ground is the overlap, integrated numerically rather than looked up, so it stays correct when you give the two groups different spreads.
Key terms
- Cohen's d
- The gap between two means expressed in pooled standard deviations.
- Pooled standard deviation
- A combined estimate of the two groups' spread. It is d's denominator, which is why changing either group's spread moves d.
- Overlap
- The share of the two distributions occupying the same range. It is a translation of d into something a reader can picture.
- Effect size and test statistic
- d describes how big a difference is; a test statistic describes how well the study was placed to detect it. They answer different questions.
Two distributions, four ways of describing the gap
d counts the gap in pooled standard deviations. The other three translate that into shares of people, which is what a reader can actually picture.
Two overlapping distributions
| Index | Value | What it counts |
|---|
Saying it without overstating it
What the sample size changes, and what it does not
| Quantity | Value | Moves with n? |
|---|
Challenge — which index is being described?
Five sentences from five fictional results paragraphs. Match each one to the quantity it is actually describing.
What this demonstrates
d is a ratio, so both parts of it matter
Cohen's d is the difference between the means divided by the pooled standard deviation. Because the denominator is in the same units as the numerator, the result has no units at all, which is what lets an eight-point gap on a fictional task be compared with a two-second gap in a reaction-time study. It also means the same raw difference produces quite different effect sizes in groups of different spread — drag one group's standard deviation and watch d move while neither mean does.
The sample size cannot touch it
d estimates a property of the populations. Collecting more people gives you a better estimate of it, not a bigger one. A t-statistic and a p-value, by contrast, grow with the square root of n without limit, which is why a study of four hundred can report an unmissable p-value for a difference nobody would notice. The disclosure under the chart puts both on screen at once.
Overlap is the honest translation
Effect sizes in standard-deviation units are hard to picture, and the conventional labels invite the picture to be much cleaner than it is. At d = 0.8, called large, 69% of the two distributions still overlaps and 71% is the chance that a randomly chosen member of the higher group beats a randomly chosen member of the lower one — which leaves 29% going the other way. Reporting the overlap alongside the effect size is the cheapest available protection against overstating a finding.
Small, medium and large were always rough
Cohen offered those labels as a stopgap for fields with no accumulated knowledge of their own, and said so. In a field with a literature, the useful comparison is with other effects in the same area, and in an applied context it is with a threshold of practical importance decided before the data arrive. A d of 0.2 can be transformative in a public-health intervention and negligible in a laboratory task, and no benchmark table can know which is which.
For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.