Research Methods
An effect size is a statement about two whole distributions, not about two averages. The picture that matters is how much of them sit on top of one another — and it is almost always more than people expect.
Simulated — illustrative distributions, no real group data
By the end you should be able to read Cohen's d as overlap between two distributions, and say why sample size cannot move it.
About 25 minutes. Nothing you do here is saved or sent anywhere.
Every index on this page describes two distributions. None of them says anything about any individual. Even at d = 1.0 — twice Cohen's "medium" — nearly a third of the two distributions still overlaps, and roughly one person in four drawn from the higher group scores below a person drawn from the lower one. Sentences of the form "group X are more Y" almost always overstate what an effect size supports.
Two fictional groups are compared on a 100-point task. Both have a standard deviation of 10; their means are 50 and 58. That is d = 0.80, which Cohen's rough labels would call a large effect.
Two normal distributions on one axis. The shaded common ground is the overlap, integrated numerically rather than looked up, so it stays correct when you give the two groups different spreads.
d counts the gap in pooled standard deviations. The other three translate that into shares of people, which is what a reader can actually picture.
| Index | Value | What it counts |
|---|
| Quantity | Value | Moves with n? |
|---|
Five sentences from five fictional results paragraphs. Match each one to the quantity it is actually describing.
Cohen's d is the difference between the means divided by the pooled standard deviation. Because the denominator is in the same units as the numerator, the result has no units at all, which is what lets an eight-point gap on a fictional task be compared with a two-second gap in a reaction-time study. It also means the same raw difference produces quite different effect sizes in groups of different spread — drag one group's standard deviation and watch d move while neither mean does.
d estimates a property of the populations. Collecting more people gives you a better estimate of it, not a bigger one. A t-statistic and a p-value, by contrast, grow with the square root of n without limit, which is why a study of four hundred can report an unmissable p-value for a difference nobody would notice. The disclosure under the chart puts both on screen at once.
Effect sizes in standard-deviation units are hard to picture, and the conventional labels invite the picture to be much cleaner than it is. At d = 0.8, called large, 69% of the two distributions still overlaps and 71% is the chance that a randomly chosen member of the higher group beats a randomly chosen member of the lower one — which leaves 29% going the other way. Reporting the overlap alongside the effect size is the cheapest available protection against overstating a finding.
Cohen offered those labels as a stopgap for fields with no accumulated knowledge of their own, and said so. In a field with a literature, the useful comparison is with other effects in the same area, and in an applied context it is with a threshold of practical importance decided before the data arrive. A d of 0.2 can be transformative in a public-health intervention and negligible in a laboratory task, and no benchmark table can know which is which.