Module 02
Research Methods
Tools for design and inference — the ideas students most often meet as formulae first and intuition second. How much a statistic moves from sample to sample, what uncertainty around an estimate actually means, and what a test statistic is responding to when it changes.
21 tools published
Tools in this module
-
Research Question to Method Mapper
Six fictional research questions, read one at a time. For each, the learner decides what kind of claim the question is trying to support, whether it is qualitative, quantitative or mixed, what design it implies, and what they would still need to know before naming an analysis. Feedback addresses each part separately and marks readings as well supported, defensible or hard to defend rather than right or wrong. No statistical test is ever named.
-
Operationalisation Laboratory
Three fictional constructs, each modelled as five facets, and six candidate operational definitions apiece drawn from self-report, behavioural-trace, observational and performance-based families. Ticking measures moves a live facet coverage map and a plain-language list of what each measure also records. Coverage rises as measures are added, and so does reactivity, while feasibility falls - and full coverage of a facet list is still not the construct.
-
Confound Detective
Three fictional studies with confident causal claims, each run in two phases. First the learner classifies five variables as confounds, nuisance variables, part of the manipulation or held constant, and receives a written note on each. Then a repair bench opens: ticking design fixes redraws a causal diagram and moves the estimated effect back towards a true effect only the tool knows. One repair in every study - recruiting more participants - is there to be seen doing nothing at all.
-
Sampling Bias Simulator
A generated population of 4,000 fictional students whose true average weekly study hours is printed on screen. Learners recruit from it by convenience, self-selection, year quotas, stratified random sampling or simple random sampling, set how much more likely commuters and students with jobs are to decline, and watch repeated estimates land on a number line beside the truth. Raising the sample size tightens the cloud of estimates without moving it, which is the whole distinction between variability and bias.
-
Thematic Analysis Coding Laboratory
Six invented interview extracts about asking for help at university, worked one at a time. The learner writes a code of their own, optionally borrows from a bank of six candidate codes that each explain what they do rather than whether they are right, and then sets their reading beside three defensible coding passes over the same words - semantic, latent and situated. Each extract closes with a reflexive prompt about what the reader brought with them. Nothing is scored, and inter-rater agreement is explicitly not the criterion.
-
Theme or Topic?
Three clusters of coded extracts from an invented dataset, each with four candidate statements. The learner sorts every candidate into topic summary, useful staging post, developed theme, or beyond what the data can carry, and receives a written note on all four. The three over-reaching candidates fail in three different ways: one asserts a cause, one attributes intent to an institution nobody interviewed, and one imports a construct together with a between-person comparison. A second challenge asks which rewrites would give a weak section a central organising concept.
-
Reflexivity and Alternative Theme Builder
The same six invented extracts used across this module's qualitative tools, paired with one of three research questions and one of three theoretical sensitivities. Each of the nine pairings yields a prepared thematic account with its themes, a map of which extracts sit at the centre and which recede, what it illuminates and what it leaves less visible. Four pairings are strong, three workable with a stated cost, and two pull against themselves. Accounts can be saved and compared side by side, and an accountability check tests five candidate claims against the reading currently selected.
-
Sampling Distribution and p-Value Simulator
Two groups of fictional participants are drawn from the same population, so the true difference is exactly zero. Running that study a thousand times builds the sampling distribution of the mean difference, and one observed difference is marked on it with the tails at least that far from zero shaded. The tail area counted in the pile is shown beside the tail area the model predicts. Sample size and variability can be changed, so the same observed difference can be watched moving from p = .5 to p below .001 without the finding changing at all.
-
Confidence Interval Laboratory
Two staged experiments. In the first, a simulated population has a true mean the studies never see, and repeated samples produce a stack of confidence intervals whose misses are marked, so coverage can be counted rather than asserted; sample size and variability change the width of those intervals and leave the coverage alone. In the second, four fictional trials are judged one at a time against a threshold of practical importance that the learner sets, which makes a precise trivial finding and a wide uninformative one look as different as they are. A select-all challenge then works through the standard misreadings of a 95% interval.
-
ANOVA F-Ratio Visualiser
Two staged experiments on the logic of the one-way F ratio. In the first, three simulated groups are drawn from populations whose separation, within-group spread and size the learner controls, and a real ANOVA on the generated sample is shown as a strip plot with each group mean's reach away from the grand mean drawn in; the point is that F responds to both halves of a fraction, and that noise enters as its square. In the second, three visibly different patterns of group means are constructed to have identical between-groups sums of squares, so all three produce F(2, 87) = 7.50 - which is what a significant omnibus test can and cannot tell you.
-
Factorial ANOVA Interaction Detective
Four editable cell means are the whole dataset of a fictional 2x2 design, and the marginal means, the two main effects, the interaction and the plot are all arithmetic performed on them. Learners predict all three effects from a table of cell means before anything is revealed, then load crossover, fan, parallel, hidden-effect and deliberately trivial patterns or drag their own. Two controls decide whether the picture means anything - the vertical scale and the uncertainty in each cell - and a fill-in-the-sentence challenge requires the interaction to be described in words rather than announced.
-
ANCOVA / MANOVA Decision Laboratory
Learners assign a role to every variable in a fictional evaluation before choosing an analysis, then meet two experiments. The first generates eighty students and shows what covariate adjustment does: with random allocation it barely moves the estimate and buys precision instead, with a baseline gap it moves it a long way, and with unequal regression slopes it stops being a single number at all - while a menu switching random allocation for intact classes changes nothing in the arithmetic and everything in the interpretation. The second plots two outcomes at once and compares each on its own with the joint Mahalanobis separation, so that a multivariate test is seen to ask a different question rather than a cheaper one. Four closing scenarios include two whose answer is neither analysis.
-
The Normal Curve and z-Scores
A normal density is drawn on fixed score and density axes, so that moving the mean slides the curve and raising the standard deviation visibly flattens it - the peak height is printed beside the picture and halves when the standard deviation doubles. The learner places a raw score, reads its z, its percentile and both tail areas, and highlights the one, two and three standard-deviation regions to find that their areas never move however far the two parameters are dragged. A second experiment puts the same raw mark on two different distributions and withholds both z-scores until a judgement has been committed.
-
Central Limit Theorem Simulator
One picture with two panels on a shared value axis keeps the three objects apart: the population above, the most recent sample drawn as ticks along its axis, and beneath them the pile of sample means. Four populations are offered - normal, strongly right-skewed, uniform and bimodal - each with exactly known mean, standard deviation, skewness and kurtosis, so the tool can print what theory predicts for the sample mean beside what the simulation produced, row by row. Because the predicted skewness of the mean is the population's divided by the square root of n, the familiar advice that thirty is enough becomes something a student can check and, for the skewed population, disprove.
-
Cohen's d and Distributional Overlap
The learner estimates on a slider what share of two distributions overlaps at a conventionally large effect before anything is revealed, and almost always guesses far too little. The explorer then draws two normal distributions with the common ground shaded and reports four descriptions of the same gap side by side: Cohen's d, the overlap coefficient, the probability of superiority and U3. Group spreads can be pulled apart, so the pooled standard deviation becomes visible arithmetic and the overlap has to be integrated rather than read off the equal-variance formula, and a sample-size slider moves the t statistic and the p-value across orders of magnitude while the effect size refuses to move at all.
-
Independent-Samples t-Test: The Null Distribution
The tool works from summary statistics, because that is all a t-test needs, and puts the whole test into one picture: the distribution of t if the two population means were equal, with the tails beyond the observed statistic shaded, the critical values as dashed rules and the standard normal drawn faintly behind so the heavier tails at small degrees of freedom can be seen rather than asserted. Two preset buttons run the same two groups at fifteen and at a hundred per group, so the same finding is read off twice with Cohen's d fixed at 0.50 and p falling from .182 to below .001. A closing challenge asks the learner to choose between four write-ups of one non-significant result.
-
Statistical Power and Type M Error
Experiment 1 draws the two sampling distributions of the observed effect size - one centred on zero, one on the effect you have assumed - with the rejection threshold between them, so alpha, beta and power become three areas that move when the assumption, the sample size or the threshold moves; a target-power control reports the sample size a planned study would need at each level. Experiment 2 runs two thousand genuinely simulated studies of one real effect and shows both halves of the Type M argument at once: the average of every estimate lands on the truth, and the average of the significant ones is nearly three times it at fifteen per group. A prediction is required before the run.
-
Correlation: Linearity, Outliers and Shared Variance
Six generated datasets on one frame, with three controls that each carry a lesson: the dataset menu shows a clean curve whose Pearson r is essentially zero, a slider drags one ringed observation towards the corner and swings r from 0.02 to 0.75 in a twelve-point sample while barely moving it in a hundred-and-twenty-point one, and a units menu multiplies y so that the slope moves by a factor of ten and r does not move at all. A disclosure recomputes everything with the ringed point removed. The challenge is an eyeball test: estimate r for three plots on sliders and find out which way your error runs.
-
Regression: Intercept, Slope and Least Squares
Thirty generated observations and a line the learner moves by hand. Every residual is drawn, optionally as the literal square whose area is being minimised, and the sum of squared errors updates live against a goal banner; a reveal then shows the least-squares solution, which no amount of dragging beats. The region outside the observed range of x is shaded, so the intercept can be seen to be a prediction for somebody who does not exist, and a disclosure shows the identical fit with x centred on its mean - same slope, same predictions, same residuals, and an intercept that finally describes a real part of the data. The challenge asks for one interpolated prediction, one extrapolated one, and a decision about which to report.
-
Homoscedasticity and Residual Diagnostics
One generated dataset drawn twice on a shared fitted-value axis - the raw scatterplot above, residuals against fitted values below - so that a variance pattern which is arguable in the first is unmistakable in the second. Five patterns of residual spread can be dialled from flat to severe, and the estimated slope keeps landing on the true 0.50 while the standard error stops being trustworthy. A disclosure reports the exact standard error the generating model implies beside the one the classical formula is aiming at, and confirms it with the coverage of both intervals across four hundred repeated samples. The challenge is four residual plots to read, one of which is not about variance at all and one of which cannot honestly be diagnosed.
-
Multiple Comparisons, FWER and Forking Paths
Experiment 1 runs a family of real two-sample t-tests in which every null is true, draws one square per test with the significant ones crossed, and sets the counted family-wise error rate beside 1 minus (1 minus alpha) to the power k; a Bonferroni option then shows the false-positive rate collapse and, in the same readout, the detection rate for the real effects collapsing with it. Experiment 2 generates ninety fictional students with no group difference on anything and offers four outcomes, three exclusion rules, three subgroups and two covariate decisions - seventy-two complete, defensible analyses of the same empty data - and lets the learner search until something reaches p below .05, then reveals how many other paths would have done the same. A closing challenge separates planned multiplicity, exploratory analysis, undisclosed selective reporting and confirmatory inference.
What this module covers
These are the topics the module was planned around. The tools above cover them; the list is kept so the intended scope stays visible and so a contributor proposing a new tool can see what is already here.
- Sampling distributions Repeatedly drawing samples from a known population and watching the distribution of the sample mean take shape — the single idea most of inferential statistics rests on.
- Confidence intervals What the interval covers across repeated samples, and what it does not license you to say about the one study in front of you.
- p-values and significance testing How a p-value responds to effect size, sample size and variability when each is varied on its own rather than together.
- Correlation and regression Building a scatterplot point by point to see how outliers, range restriction and non-linearity move a coefficient that is usually reported as a single settled number.
- Statistical power The trade-off between sample size, effect size and the chance of detecting an effect that is really there.
Building one of these
All five topics above are open. If you teach research methods and have a demonstration you already run by hand, turning it into a page here makes it reusable by everyone else. The contributing guide covers the folder layout, the shared interactive shell and the accessibility checks a tool needs to pass before it is merged.