Research Methods · Simplified
Sampling Bias Simulator
Four thousand students whose true average you can see. Recruit from them, repeatedly, and watch where the estimates land.
The simulator
How you recruit decides where the estimates land
The population is generated once and never changes, so its true mean weekly study hours is known. Every sample is a real draw from it. Draw a few, then draw a lot.
Key terms
- Bias
- The estimates centre on the wrong value. Averaging more of them, or drawing bigger samples, does not move them back.
- Sampling variability
- Estimates scatter around wherever they centre, because each sample is a different set of people. Bigger samples scatter less.
- Convenience sample
- Whoever is easy to reach. Who is easy to reach is usually related to what is being measured.
- Quota sample
- Fixed numbers recruited in each category, filled however is convenient. It corrects the composition of the variable you set a quota on, and nothing else.
The solid line is the population mean, which a perfect census would return. Each bar counts studies whose sample mean landed there.
What this shows
Scatter and being wrong are different failures
A simple random sample of 20 is all over the place and centred on the truth. A convenience sample of 600 is tightly clustered and centred somewhere else. Only one of those two problems is fixed by recruiting more people.
Key idea: Sample size buys precision. Nothing about it buys the right answer. If who ends up in your sample is related to what you are measuring, the estimates centre on the wrong value and go on centring there however many you collect. A quota fixes the composition of the variable you set a quota on and leaves every other selection problem exactly where it was, which is why a table showing a perfect year breakdown is not evidence that a sample is any good.
The population is generated and the selection weights are invented, chosen to make the effect legible rather than to estimate how much more likely a commuter is to be in a library on a Tuesday. Real non-response depends on the topic, the mode, the incentive and the season. Bias in a mean is also not the only bias selection produces: it distorts variances, subgroup comparisons and associations too, and not always in the same direction.
The longer version adds five recruitment methods including stratified sampling, a non-response dial that applies to every method, a composition table and a select-all challenge. It is at Sampling Bias Simulator in the main collection.