Research Methods
Run enough tests on nothing and something will come out significant. The uncomfortable part is that you do not have to run them deliberately, or even know that you did.
Simulated — generated data containing no effect at all
By the end you should be able to say what family a family-wise error rate belongs to, and how analytic flexibility inflates false positives without anyone cheating.
About 30 minutes. Nothing you do here is saved or sent anywhere.
The second experiment lets you search a fictional dataset until something comes out significant. It is there to show you what that search does to the meaning of the number you end up with, and every choice it offers is one a careful researcher might make in good faith. Nobody in this laboratory is cheating. That is exactly why the problem is hard, and why the answer is transparency about what was decided when, rather than better intentions.
A fictional study runs 20 independent tests at α = .05. Every single null hypothesis is true: there is nothing to find anywhere in the data.
Each square below is one test in one simulated experiment. A square is marked when that test came out significant. Most of the tests have nothing to find; you can give a few of them a real effect and watch what a correction does to those as well.
The family-wise error rate is a property of the whole family, not of any one test in it. Every test in the family is behaving exactly as advertised.
| In the most recent experiment | Count |
|---|
| Quantity | Predicted | Simulated |
|---|
One fictional dataset: ninety students, randomly assigned to two versions of a revision task, measured on four outcomes, with a baseline test, a response-speed measure and a cohort label. The data were generated with no difference between the two groups on anything. Below are the choices a careful analyst might reasonably make. Make them, and see what happens.
Four outcomes, three exclusion rules, three subgroups and two covariate decisions. Nobody would call any single one of those choices dishonest.
| Quantity | Value |
|---|
| Range of p | Analyses |
|---|
Multiplicity is not one thing. Label each of these four fictional studies. Only one of the four is a problem, and it is not the one with the most tests in it.
Each test keeps its own 5% error rate; what climbs is the chance that somewhere in the family a null result is declared significant. Twenty independent tests at .05 give roughly a two-in-three chance of at least one false positive, with nothing wrong in the data. Analytic flexibility does the same damage without any test being repeated, because the path was chosen after seeing where the result fell.