Research Methods · Simplified
Confidence Interval Laboratory
The truth stays still. The intervals move around it. Ninety-five per cent is a statement about the intervals.
The simulator
One truth, many studies, one interval each
A simulated population whose true mean change in walking is exactly 12 minutes a day. Each study draws its own sample and reports its own interval. Only one thing on this page changes how often those intervals contain the truth.
Key terms
- Confidence interval
- A range computed from one sample by a procedure that, repeated, captures the true value a stated share of the time.
- Coverage
- The share of intervals that actually contain the true value. It is a property of the procedure across many studies, not of any one interval.
- Confidence level
- The share you asked for when you chose the multiplier: 80, 90, 95 or 99 per cent.
- Width
- How long the interval is. Sample size and population spread change the width. Neither changes the coverage.
One row per study, most recent at the top. The solid line is the true value. Intervals that miss it are drawn as a dashed line.
What this shows
Width and coverage are different things
Sample size and population spread change how long the intervals are. Neither changes how often they contain the truth. Only the confidence level does, and it does so by changing the width in exchange.
Key idea: A confidence level is a property of the procedure, not of the interval in front of you. The one interval your study reports either contains the true value or it does not, and you will never know which. What the ninety-five per cent describes is the long run: repeat the procedure and that share of the intervals will capture the truth. That is why "there is a 95 per cent chance the true value is in this interval" is the wrong sentence, even though the right sentence is harder to say.
The population standard deviation is treated as known here, because the simulation states it, which is why the multiplier is a normal quantile and why every interval in a run has exactly the same width. A real study estimates the spread from its own sample and uses a t multiplier, which widens the interval at small samples and makes the widths vary from study to study. That estimation is a different tool. Coverage is also exact only because the model is exactly right here; real coverage degrades when the assumptions do.
The longer version adds a second experiment in which four fictional trials are judged against a smallest-change-that-matters threshold, separating an interval that is informative from one that is merely narrow, plus a select-all challenge. It is at Confidence Interval Laboratory in the main collection.