Research Methods · Simplified
What Gets Published From a Small Study
Every study here is honest, unbiased and correctly analysed. The ones that reach significance still overstate the effect, and nobody had to do anything wrong.
The simulation
Two thousand studies of the same real effect
The effect below is genuinely there, in every single one of these studies. Each research team draws its own sample, runs a correct t-test, and reports what it finds. Only the significant ones tend to get written up.
Key terms
- Power
- The proportion of studies that reach significance when the effect really is there. It is a property of the study design, not of the result.
- Type M error
- An error of magnitude. How much a significant finding overstates the true effect, on average.
- Type S error
- An error of sign. A significant finding that points in the opposite direction to the truth.
- Exaggeration ratio
- The average effect among significant studies, divided by the true effect. A ratio of 2 means the published literature reports the effect as twice its real size.
One bar per slice of the range. The solid part of each bar is the studies that reached significance; the outlined part is the studies that did not. The dashed line is the true effect and the solid line is the average of the significant studies alone.
What this shows
Filtering on significance is what does the damage
No individual study here is biased. The average across all two thousand of them lands on the true effect. The bias appears the moment you look only at the ones that reached significance.
Key idea: To be significant, a small study has to produce a large estimate: that is what the threshold means. So when power is low, the significant studies are precisely the ones that got a lucky, unusually large sample, and the published effect is systematically too big. Raising the sample size fixes this, not by making anyone more careful, but by shrinking the gap between the effect a study needs in order to be significant and the effect that is actually there. This is also why a significant result from a small study is weak evidence about the size of an effect even when it is correct about its existence, and why replications so reliably come out smaller than the original.
Three things this does not show. It is not about fraud or p-hacking: every simulated team here does everything correctly, and the exaggeration arises purely from which results get selected. It is not an argument that significant findings are usually wrong; at decent power the exaggeration is small and the sign errors vanish. And computing power after the fact, from the effect a study happened to observe, tells you nothing that the p-value did not already tell you, because the observed effect is exactly the thing that has been distorted. Power is something to work out before running a study, from the smallest effect that would matter.
The longer version adds alpha, beta and power as three moving areas on two sampling distributions, a sample-size calculator for a target power, and a worked treatment of why observed power adds nothing. It is at Statistical Power and Type M Error in the main collection.