Research Methods
Multiple Comparisons, FWER and Forking Paths
Run enough tests on nothing and something will come out significant. The uncomfortable part is that you do not have to run them deliberately, or even know that you did.
Simulated — generated data containing no effect at all
Learning objective
By the end you should be able to say what family a family-wise error rate belongs to, and how analytic flexibility inflates false positives without anyone cheating.
About 30 minutes. Nothing you do here is saved or sent anywhere.
This is not a technique
The second experiment lets you search a fictional dataset until something comes out significant. It is there to show you what that search does to the meaning of the number you end up with, and every choice it offers is one a careful researcher might make in good faith. Nobody in this laboratory is cheating. That is exactly why the problem is hard, and why the answer is transparency about what was decided when, rather than better intentions.
- Answer the prediction question to open Experiment 1.
- Run a family of tests where every null is true, and count the hits.
- Apply a Bonferroni correction and read both of its consequences.
- In Experiment 2, analyse one empty dataset until something works.
First, a prediction
A fictional study runs 20 independent tests at α = .05. Every single null hypothesis is true: there is nothing to find anywhere in the data.
Experiment 1 — the family of tests
Each square below is one test in one simulated experiment. A square is marked when that test came out significant. Most of the tests have nothing to find; you can give a few of them a real effect and watch what a correction does to those as well.
Key terms
- Family
- The set of tests a claim is being made across. The error rate here is a property of that set, not of any test inside it.
- Family-wise error rate
- The probability of at least one false positive somewhere in the family, when every null in it is true.
- Bonferroni correction
- Dividing α by the number of tests. It has a second consequence as well as the intended one.
- Forking paths
- The many analysis decisions available, each defensible on its own, which between them produce many possible results from one dataset.
One experiment, k tests
The family-wise error rate is a property of the whole family, not of any one test in it. Every test in the family is behaving exactly as advertised.
Nothing run yet
| In the most recent experiment | Count |
|---|
Reading this correctly
Theory beside simulation
| Quantity | Predicted | Simulated |
|---|
Experiment 2 — the garden of forking paths
One fictional dataset: ninety students, randomly assigned to two versions of a revision task, measured on four outcomes, with a baseline test, a response-speed measure and a cohort label. The data were generated with no difference between the two groups on anything. Below are the choices a careful analyst might reasonably make. Make them, and see what happens.
Seventy-two defensible analyses of nothing
Four outcomes, three exclusion rules, three subgroups and two covariate decisions. Nobody would call any single one of those choices dishonest.
No analyses run yet
| Quantity | Value |
|---|
Every path you could have taken
| Range of p | Analyses |
|---|
Challenge — four studies, four practices
Multiplicity is not one thing. Label each of these four fictional studies. Only one of the four is a problem, and it is not the one with the most tests in it.
What this demonstrates
Each test keeps its own 5% error rate; what climbs is the chance that somewhere in the family a null result is declared significant. Twenty independent tests at .05 give roughly a two-in-three chance of at least one false positive, with nothing wrong in the data. Analytic flexibility does the same damage without any test being repeated, because the path was chosen after seeing where the result fell.
For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.