Research Methods
Three fictional studies, each with a confident causal claim. Sort the variables that travel with the manipulation from the ones that only add noise — then repair the design and watch the estimate move back towards the truth.
Simulated — fictional studies and a transparent bias model
By the end you should be able to state the two conditions a variable must meet to confound an effect, and test a candidate against both.
About 30 minutes. Nothing you do here is saved or sent anywhere.
Each study has a true effect that only the tool knows, plus a bias contributed by each unresolved confound. Those figures are chosen to make the argument visible. They are not estimates of how large real confounding is in any real study, and no real trial, ward, module or dataset is described.
A study reports that students who use the library more get higher marks. A critic objects that motivation is a confound.
Each study runs in two phases. First you classify the variables; then the repair bench opens and you fix the design. The study card stays at the top of the stage throughout.
A variable can only bias a comparison if it travels with the conditions and touches the outcome. Everything else is noise, treatment, or beside the point.
| Confound | Status | Bias it adds |
|---|
A randomised trial allocates 200 volunteers to a six-week exercise programme or a waiting list, and measures depressive symptoms before and after. The exercise group improves more.
It must differ systematically between the conditions, and it must plausibly touch the outcome. A variable with only the first is irrelevant; a variable with only the second adds noise and widens your interval without moving your estimate. Students who learn to list "extra variables" find confounds everywhere and cannot say which ones matter.
Noise makes an estimate imprecise, and more data fixes it. Bias makes an estimate wrong, and more data makes the wrong answer more confident. The repair bench contains a sample-size button precisely so that it can be seen doing nothing.
Randomising breaks the link between the condition and everything else, measured or not, known or not. Adjusting for a covariate handles that one variable, provided it was measured, measured well, and entered into the model in roughly the right form. That is a much longer list of things to be right about, which is why the two are not interchangeable however similar the output looks.
The spaced-repetition schedule in study 1 is not a confound; it is what the app is. Removing it does not clean up the comparison, it deletes the manipulation. Deciding which features are the active ingredient is a theoretical question, and answering it needs a condition designed to isolate it rather than a statistical control.