Research Methods

Confidence Interval Laboratory

The truth stays exactly where it is. The intervals jump about. Count how many of them catch it — then work out whether catching it is the same thing as saying something worth knowing.

Simulated — a generated population and four fictional trials

Learning objective

By the end you should be able to say what the 95% in a confidence interval is a property of, and why it is not a claim about the interval on screen.

About 25 minutes. Nothing you do here is saved or sent anywhere.

The sentence to be careful with

"There is a 95% probability that the true value lies inside this interval" is the phrasing almost everyone reaches for, and it is not what the method delivers. The 95% describes how often the procedure catches the truth when the whole study is repeated. Any single interval either contains the true value or it does not, and in a real study nobody knows which. This laboratory can show you which only because it invented the truth first.

  1. Answer the prediction question to open Experiment 1.
  2. Draw studies one at a time, then a hundred at once, and count the misses.
  3. Change the sample size and confidence level, and watch which number moves and which does not.
  4. In Experiment 2, set what counts as a difference worth having, then judge four fictional findings against it.

First, a prediction

One hundred fictional research teams each run the same study on the same population and each reports a 95% confidence interval for the population mean. Because the population is simulated, we know the true mean exactly, so we can check every interval.

How many of the 100 intervals will contain the true population mean?

Challenge — what the interval says

A fictional study of 60 people reports a mean change of 4.8 minutes of walking per day, 95% confidence interval [1.2, 8.4].

Which statements are correct? Select all that apply — more than one is.

What this demonstrates

The 95% belongs to the method, not to your interval

Nothing in the calculation is random once the sample has been collected: the limits are fixed numbers, and the population mean is a fixed number, so the interval either contains it or it does not. What repeats is the procedure. Run it a hundred times and about ninety-five of the intervals it produces will contain the truth. That is a guarantee about the long run. It says nothing about which of the hundred you happen to be holding.

Width and coverage are different properties

Sample size and variability change how wide the intervals are. They do not change how often the intervals catch the truth: a 95% procedure covers about 95% of the time whether the intervals are enormous or tiny. The confidence level is the only control that moves coverage, and it buys that coverage with width. A 99% interval catches the truth more often precisely because it is less informative.

Excluding zero is a low bar

An interval that excludes zero tells you the data are hard to reconcile with exactly no difference. That is a weaker claim than it sounds, because exactly no difference is rarely the interesting hypothesis. The second fictional trial excludes zero with great precision and reports a change nobody would install an app for. Deciding what would count as worth having is a judgement made before the data arrive, and it is what turns an interval into a conclusion.

A wide interval is not a negative finding

The third trial's interval runs from −6 to +34. It is routinely written up as "no significant effect", and it is compatible with no effect, with a small effect and with an effect far larger than anyone predicted. The correct description is that the study was too small to distinguish between them. Reporting the interval makes that obvious; reporting only whether it crossed zero hides it.

Download activity HTML

For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.