Research Methods · Simplified
The Garden of Forking Paths
One dataset, several perfectly defensible choices, and a p-value at the end of each.
The analysis
Your study, and the decisions nobody wrote down
Ninety students were assigned to a new revision method or to the usual one. You measured four things afterwards, recorded a baseline test score and how quickly each student worked, and you know which are first-year students.
Now analyse it. Every choice below is one a careful researcher might make in good faith, and none of them is cheating.
Key terms
- p-value
- How often a difference at least this large would turn up if there were really no difference, in an analysis decided before the data were seen.
- Unadjusted
- The two groups are compared on the outcome exactly as measured, with nothing taken out of it first.
- Adjusted for baseline score
- The part of the outcome that the baseline test already predicted is removed before the groups are compared. The baseline is acting as a covariate: a variable held constant so the comparison is not confused by it.
- Exclusion rule
- A stated reason for leaving some participants out before analysing, such as working implausibly fast. Which rule you pick changes who is in the sample.
- Subgroup
- A part of the sample analysed on its own, such as first-year students. The comparison is then about those people only.
- Significant at 0.05
- Shorthand for a p-value below 0.05. It is a conventional threshold, not a discovery, and it says nothing about how large or important a difference is.
No analysis run yet.
What you have run
Try a few. Change one choice at a time, or change your mind entirely: nothing here records what you looked at first.
Every path
There was nothing there
One dot per complete, reportable analysis. The shaded region on the left is everything that would have been called significant.
Key idea: A p-value answers a question about a procedure: how often would data like this arise if there were nothing to find, under this analysis, specified in advance. Choose the analysis after seeing the data and the procedure is no longer the one the p-value describes. Nobody here ran seventy-two tests. A researcher walks one path, reports one number, and the other seventy-one never existed on paper. That is what makes this hard to catch: the published analysis looks exactly like a pre-specified one.
This is not a demonstration of dishonesty, and none of the choices offered is a bad one. Analysts face exactly these decisions and resolve them with reasons. The problem is that the reasons are available after the data are, and a decision that would have gone the other way on other data is a decision the p-value has not accounted for. Note also that these paths share participants and outcomes, so they are heavily correlated: this is why no simple correction applies to them, and why pre-registration and reporting the whole analysis space are the usual answers rather than dividing by seventy-two. The dataset is invented and the exact count of significant paths is a property of this one simulated sample.
This is the forking-paths half of a longer activity. The full version adds a simulation of the family-wise error rate across a family of independent tests, a Bonferroni correction shown together with its cost in detection, and a set of vignettes to classify. It is at Multiple Comparisons, FWER and p-Hacking in the main collection.