Research Methods

Statistical Power and Type M Error

Power is a plan, not a post-mortem: it is worked out before the study, for an effect you have assumed. And when power is low, the results that pass the filter are not just fewer — they are bigger than the truth.

Simulated — generated studies from a documented seed

Learning objective

By the end you should be able to say what power is planned for, and why a filtered literature exaggerates the effects it reports.

About 30 minutes. Nothing you do here is saved or sent anywhere.

Low power does not bias your estimate

This is the part that is easy to get wrong in the other direction. An underpowered study's effect-size estimate is unbiased: run a thousand of them and their average lands on the truth. The exaggeration appears only when you look at the significant ones, because a small study can only reach significance by overshooting. It is a property of the filter, not of the studies.

  1. Answer the prediction question to open Experiment 1.
  2. Move the true effect, the sample size and α, and watch the four areas.
  3. Use the target-power control to find the sample size a study would need.
  4. In Experiment 2, commit to a prediction, then run two thousand studies.

First, a prediction

A fictional study compares two groups of 20 people each, using a two-tailed test at α = .05. The true effect in the population is d = 0.50 — Cohen's "medium".

What is the probability that this study returns a significant result?

Challenge — what power is and is not

Which statements are correct? Select all that apply — four of the seven are.

What this demonstrates

Power is four areas and a decision

Draw what the observed effect size does when the truth is zero, draw what it does when the truth is what you assumed, and put a threshold between them. The tail of the first curve beyond the threshold is α. The body of the second curve on the wrong side of it is β. Power is the rest. Everything you can do to a study — more participants, a cleaner measure, a larger manipulation, a looser threshold — works by moving those curves apart or narrowing them.

The sample size you need grows as the square of the reciprocal

Halving the effect size you want to detect roughly quadruples the sample you need. At d = 0.5 the classic 80% target costs 64 per group; at d = 0.25 it costs 253. That single fact explains most of what has gone wrong in fields where small effects are studied with small samples, and no amount of analysis after the fact recovers it.

Observed power is not a diagnosis

Power computed after a study, using the effect the study itself observed, is a one-to-one function of the p-value: a high p always gives low observed power, by construction. It therefore adds no information at all, and "the analysis was underpowered, as shown by the observed power of 21%" is a restatement of the p-value rather than an explanation of it. Power calculations belong in a protocol, not in a discussion section.

The exaggeration is in the filter, not in the studies

This is the part worth being careful about. Every simulated study in Experiment 2 estimates the effect without bias, and the average of all two thousand estimates lands on the truth. But a small study can only clear the significance threshold by overshooting, so the subset that clears it has a much larger average. Publish only those and the literature reports an effect two or three times the real one, with nobody having done anything wrong at any single step. At very low power some of those significant results also point in the wrong direction entirely.