Research Methods
F is a fraction. Pull the group means apart and the top grows; make the groups noisier and the bottom grows. Nothing about the test makes sense until you can see both halves of it at once.
Simulated — generated scores from a documented seed
By the end you should be able to read F as a ratio of two variances, and say what question the omnibus test actually answers.
About 25 minutes. Nothing you do here is saved or sent anywhere.
One-way ANOVA asks a single question: are all the population means equal? It is not three t-tests wearing a coat, and a small p does not identify which groups differ, in which direction, or by how much. Those are separate questions with separate answers, and the second experiment on this page is built to make that unavoidable.
Three fictional groups of 30 people each sit the same 100-point test after a different revision method. In the first study the population means are 44, 52 and 60, and the within-group standard deviation is 6. A second study has exactly the same three population means, but the scores within each group are twice as spread out.
Three groups, labelled A, B and C so that nothing about the story does the explaining. Each dot is one simulated score, each heavy bar is a group mean, and the dashed line is the grand mean.
The distance of each bar from the dashed line is the between-groups part. The scatter of dots around their own bar is the within-groups part. F is one divided by the other, after each is turned into a mean square.
| Group | n | Mean | SD |
|---|
| Source | SS | df | MS |
|---|
A fictional study of three revision methods reports F(2, 87) = 7.50, p = .001, with 30 people per group and a within-group standard deviation of 6. Below are three patterns of group means. Look at each one, then say which of them produced that F.
Browse the patterns freely. The F each one would produce stays hidden until you commit to an answer.
| Group | Mean | Distance from grand mean |
|---|
| Pattern | Means | SS between | F(2, 87) |
|---|
A fictional study of three revision methods, 30 people per group, reports F(2, 87) = 7.50, p = .001.
If the population means really are equal, the scatter of the group means around the grand mean is just sampling noise. Both mean squares then estimate the same population variance. Their ratio then wanders around 1. When the population means are not equal, the numerator picks up something extra that the denominator never sees, and the ratio climbs. That is the entire logic, and it is why a ratio is used rather than a difference.
Separating the means raises it. Increasing the within-group spread lowers it, and does so as the square. Double the within-group standard deviation and the amount by which F exceeds 1 divides by four. Increasing n raises it, because between-groups sum of squares grows with n while the within-groups mean square does not. A large F is therefore not a statement about how big the difference between the methods is.
A significant F says the data sit awkwardly with the hypothesis that all the population means are equal. It does not say which means differ, in which direction, or by how much. Experiment 2 makes this concrete: three genuinely different patterns of means — one evenly spaced, one with a high outlier, one with a low outlier — produce exactly the same F. Anything you want to say about particular groups needs a planned contrast or a post-hoc comparison, and those carry their own multiplicity problem.
Running every pairwise t-test would answer three separate questions and would inflate the chance of at least one false positive, which is the reason the omnibus test exists. But the relationship runs the other way too: with only two groups, F is exactly the square of the independent-samples t. They are the same test written twice, and the generalisation to three groups is where they part company.