Research Methods · Simplified
Sampling Distribution and p-Value Simulator
A world in which nothing is going on. Run studies in it, and see how far apart two groups drift anyway.
The simulator
Two groups from the same population
Both groups are drawn from one population, so the true difference is exactly zero. Every study still finds some difference. The pile of those differences is the sampling distribution, and the p-value is read off it.
Key terms
- Null model
- The assumed state of affairs the calculation is done under: here, that both groups come from the same population and the true difference is zero.
- Sampling distribution
- The spread of results you would get by repeating the same study many times when the null model is true.
- Standard error
- How much the difference between two group means moves from study to study. Here it is the population standard deviation times the square root of two over the group size.
- p-value
- The share of studies under the null model that land at least as far from zero as the one observed, counting both directions.
One mark per simulated study. The shaded regions are everything at least as far from zero as the observed difference, in either direction.
What a p-value is not
A tail area, and it points backwards
The two figures agree because they are the same quantity computed twice: once as an area under an assumed distribution, once as a count of simulated studies. The tail area is not a convention. It is a proportion of something.
Key idea: Everything on this page was computed assuming the null model, independent sampling, a pre-specified analysis and a single comparison. A p-value answers one question: if there were nothing going on, how often would a difference this far from zero turn up. It is not the probability that the null hypothesis is true, not the probability that the result happened by chance, not a measure of the size or importance of an effect, and not evidence of no effect when it is large.
No threshold is named anywhere in this activity, deliberately. Whether a tail area of a particular size should change anybody's behaviour is a decision rather than a statistical fact. Note also that the population standard deviation is treated as known here, because the null model states it, which is why the reference curve is normal. A real t-test estimates it from the sample and pays for that with heavier tails and a larger p-value at small samples, which is a different tool.
The longer version adds a three-part prediction before the simulation, progressive disclosure of the reference distribution, and a select-all challenge on interpretation. It is at Sampling Distribution and p-Value Simulator in the main collection.