Research Methods
Power is a plan, not a post-mortem: it is worked out before the study, for an effect you have assumed. And when power is low, the results that pass the filter are not just fewer — they are bigger than the truth.
Simulated — generated studies from a documented seed
By the end you should be able to say what power is planned for, and why a filtered literature exaggerates the effects it reports.
About 30 minutes. Nothing you do here is saved or sent anywhere.
This is the part that is easy to get wrong in the other direction. An underpowered study's effect-size estimate is unbiased: run a thousand of them and their average lands on the truth. The exaggeration appears only when you look at the significant ones, because a small study can only reach significance by overshooting. It is a property of the filter, not of the studies.
A fictional study compares two groups of 20 people each, using a two-tailed test at α = .05. The true effect in the population is d = 0.50 — Cohen's "medium".
The horizontal axis is the effect size a study would observe. The left curve is what that observation does when the true effect is zero; the right curve is what it does when the true effect is what you have assumed. Everything else is an area.
α is the shaded tail under the left curve; β is the shaded body of the right curve on the wrong side of the threshold; power is everything else under the right curve.
| Quantity | Value |
|---|
| Target power | Per group | Total |
|---|
Two thousand fictional studies, all of them investigating the same real effect of d = 0.35, all tested at α = .05. Each one is a genuine simulated experiment: two groups are generated, a t-test is run, an effect size is estimated. Commit to a prediction before running them.
Every study is unbiased. The literature built from the significant ones is not, and the smaller the studies the worse it gets.
| Quantity | Value |
|---|
Draw what the observed effect size does when the truth is zero, draw what it does when the truth is what you assumed, and put a threshold between them. The tail of the first curve beyond the threshold is α. The body of the second curve on the wrong side of it is β. Power is the rest. Everything you can do to a study — more participants, a cleaner measure, a larger manipulation, a looser threshold — works by moving those curves apart or narrowing them.
Halving the effect size you want to detect roughly quadruples the sample you need. At d = 0.5 the classic 80% target costs 64 per group; at d = 0.25 it costs 253. That single fact explains most of what has gone wrong in fields where small effects are studied with small samples, and no amount of analysis after the fact recovers it.
Power computed after a study, using the effect the study itself observed, is a one-to-one function of the p-value: a high p always gives low observed power, by construction. It therefore adds no information at all, and "the analysis was underpowered, as shown by the observed power of 21%" is a restatement of the p-value rather than an explanation of it. Power calculations belong in a protocol, not in a discussion section.
This is the part worth being careful about. Every simulated study in Experiment 2 estimates the effect without bias, and the average of all two thousand estimates lands on the truth. But a small study can only clear the significance threshold by overshooting, so the subset that clears it has a much larger average. Publish only those and the literature reports an effect two or three times the real one, with nobody having done anything wrong at any single step. At very low power some of those significant results also point in the wrong direction entirely.