Research Methods
Sampling Bias Simulator
A fictional population of 4,000 students whose true average you are allowed to see. Recruit from it five different ways, decide who declines, and watch the difference between an estimate that scatters and one that is systematically wrong.
Simulated — a generated population, drawn from a documented seed
Learning objective
By the end you should be able to separate sampling variability from selection bias, and say why a larger sample fixes only one of them.
About 25 minutes. Nothing you do here is saved or sent anywhere.
You can see the truth; a researcher cannot
The population mean is printed on screen because the tool generated the population. Every argument in the debrief turns on the fact that a real study has no such line to check itself against, and that a biased estimate looks exactly like an unbiased one from the inside.
- Answer the opening question to unlock the simulator.
- Pick a recruitment method and press Recruit a sample.
- Press it again several times and watch where the dots land.
- Now raise the sample size, and watch what does and does not change.
First, a prediction
A researcher wants the average weekly independent study hours of students at one university. They post a link on the university's social-media accounts and get 900 responses.
The simulator
The line marked population mean is the truth. Each dot is one recruitment exercise. Where the dots sit relative to that line is bias; how far they spread from each other is variability.
Key terms
- Sampling variability
- How far an estimate moves from one sample to the next simply because a different set of people was drawn.
- Selection bias
- A systematic offset in an estimate, because who ends up in the sample is related to what is being measured.
- Probability sampling
- A procedure in which every member of the population has a known, non-zero chance of selection. It describes the process, not the sample you end up holding.
- Non-response
- Selected people who do not take part. It matters when whether somebody takes part is related to the outcome.
- Quota
- Forcing the sample to match the population on a chosen characteristic. It acts on that characteristic and no other.
4,000 fictional students
Some commute, some hold jobs, some are in their first year. All of those things are related to how much they study, which is what makes recruitment method matter.
Where the estimates land
What this recruitment is doing
Who ended up in the sample
| Group | Population | Your sample | Average hours in the population |
|---|
Every estimate you have drawn
| Draw | Estimate | Difference from truth |
|---|
Challenge — which changes reduce the bias?
A survey of study hours has been run by advertising in the library and has produced an estimate that is 2.6 hours too high. The team has budget for exactly one change.
What this demonstrates
Scatter and offset are different failures
Sampling variability is why two honest studies of the same population report different numbers: each drew a different set of people, and the estimate wobbles around what it is aiming at. Selection bias is why a whole series of studies can wobble around the wrong thing. Only the first shrinks when the sample grows.
A large biased sample is worse than a small one
Not because the estimate is further from the truth; it is not. The interval around it is narrower and the p-value smaller, so the wrong answer is asserted with more confidence and is harder to argue with. This is precisely what happened to some famous election polls with sample sizes in the millions.
Probability sampling is a claim about the whole process
Randomly selecting names from a register makes the estimate unbiased only if the people selected actually respond. Raise the non-response tilt and the random sample drifts as far off as the convenience sample. The randomness protects the selection step; it does nothing about the step where people decide whether to answer.
Quotas fix what you measured and nothing else
Filling year-of-study quotas by convenience produces a sample with exactly the right year composition and the wrong answer, because commuting and part-time work were never balanced. Every weighting and quota scheme has this shape: it corrects the variables you thought of, and leaves the ones you did not.
For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.