Research Methods
The truth stays exactly where it is. The intervals jump about. Count how many of them catch it — then work out whether catching it is the same thing as saying something worth knowing.
Simulated — a generated population and four fictional trials
By the end you should be able to say what the 95% in a confidence interval is a property of, and why it is not a claim about the interval on screen.
About 25 minutes. Nothing you do here is saved or sent anywhere.
"There is a 95% probability that the true value lies inside this interval" is the phrasing almost everyone reaches for, and it is not what the method delivers. The 95% describes how often the procedure catches the truth when the whole study is repeated. Any single interval either contains the true value or it does not, and in a real study nobody knows which. This laboratory can show you which only because it invented the truth first.
One hundred fictional research teams each run the same study on the same population and each reports a 95% confidence interval for the population mean. Because the population is simulated, we know the true mean exactly, so we can check every interval.
A fictional walking-reminder app is tested on a simulated population whose true mean change is exactly 12.0 minutes of walking per day. Every study drawn below samples from that same population and reports an interval. The truth never moves; the intervals do.
Each horizontal bar is one study's confidence interval. The vertical rule is the true population mean, which the studies themselves never get to see.
| Outcome | Intervals | Share |
|---|
| Study | Sample mean | Interval | Contained 12.0? |
|---|
Four fictional trials of four different walking apps, each reporting a mean change in minutes walked per day with a 95% confidence interval. Before you judge them, you have to say what would count as a change worth having. That number does not come out of the data.
Set your threshold, then decide what this interval does and does not rule out.
| Quantity | Minutes per day |
|---|
| Trial | n | Reported change | 95% interval |
|---|
A fictional study of 60 people reports a mean change of 4.8 minutes of walking per day, 95% confidence interval [1.2, 8.4].
Nothing in the calculation is random once the sample has been collected: the limits are fixed numbers, and the population mean is a fixed number, so the interval either contains it or it does not. What repeats is the procedure. Run it a hundred times and about ninety-five of the intervals it produces will contain the truth. That is a guarantee about the long run. It says nothing about which of the hundred you happen to be holding.
Sample size and variability change how wide the intervals are. They do not change how often the intervals catch the truth: a 95% procedure covers about 95% of the time whether the intervals are enormous or tiny. The confidence level is the only control that moves coverage, and it buys that coverage with width. A 99% interval catches the truth more often precisely because it is less informative.
An interval that excludes zero tells you the data are hard to reconcile with exactly no difference. That is a weaker claim than it sounds, because exactly no difference is rarely the interesting hypothesis. The second fictional trial excludes zero with great precision and reports a change nobody would install an app for. Deciding what would count as worth having is a judgement made before the data arrive, and it is what turns an interval into a conclusion.
The third trial's interval runs from −6 to +34. It is routinely written up as "no significant effect", and it is compatible with no effect, with a small effect and with an effect far larger than anyone predicted. The correct description is that the study was too small to distinguish between them. Reporting the interval makes that obvious; reporting only whether it crossed zero hides it.