Personality & Individual Differences · Simplified
The Same Difference, Twice Over
Two groups differ on a questionnaire. That sentence has two completely different meanings, and the totals cannot tell you which one you have.
About the two groups
About the two groups. They are called group A and group B and have no nationality, ethnicity, language, region or culture. Where this page discusses why an item might behave differently, it does so in terms of a setting, whether disagreeing openly is invited or costly, and not in terms of a people. Nothing here is a claim about any real group, and nothing could be: every number is one you set with a slider.
Step 1 of 2
First, a judgement
A four-item assertiveness scale is given to two groups. Group B averages 0.20 points higher on the 1 to 5 scale, and with a sample of several thousand the difference is comfortably significant. What does that establish?
Step 2 of 2
Two ways to produce one number
The observed difference stays at 0.20 throughout. The slider changes only where it comes from: at one end the two groups are identical on the trait and a single item is easier for group B to endorse, and at the other end every item behaves identically and group B genuinely stands higher.
Key terms
- Latent variable
- The standing on the trait that the model posits. It is not any observed score.
- Loading
- How strongly an item reflects the latent variable. Equal loadings across groups mean a unit of the trait buys the same amount of item in both.
- Intercept
- The answer expected on an item from somebody at the middle of the trait. Two groups can differ here while their latent means match, which is what makes an item easier or harder to endorse at the same level of the trait.
- Scalar invariance
- Equal loadings and equal intercepts across groups. It is the condition that licenses comparing means, and it is a thing to be tested rather than assumed.
One bar per item, showing how much higher group B scores on it. The four bars always add to the same total. Only their shape changes, and that shape is what a test of invariance looks at.
| Item | Loading | Intercept, group A | Intercept, group B | Behaves the same? |
|---|
What this shows
The total cannot tell you which situation you are in
Key idea: At every position of the slider group B averages 0.20 points above group A, and at every position that sentence means something different. At one end there is no difference in assertiveness at all and one item is simply easier to endorse where speaking up is invited rather than costly. At the other end every item works identically in both groups and group B really does stand higher. The total score is the same number in both cases, so no amount of looking at totals, and no p-value computed from them, can separate them. What separates them is the shape of the bars, which is to say the item-level parameters, which is what a test of measurement invariance examines. That is why the test comes before the comparison rather than after it. Without it, an observed mean difference is not yet a finding about people; it is a finding about numbers that people produced.
Every figure here is a parameter that was set rather than estimated, with no sampling error, so nothing carries the uncertainty a real invariance test works with. Real tests compare the fit of nested models and reach conclusions with degrees of confidence, not certainties. Three further cautions. Failing scalar invariance does not mean a scale is worthless or that a comparison is forbidden forever: it means the comparison of raw means is not licensed, and partial invariance, where most items behave alike and a few do not, is common and can often still support a comparison. An item behaving differently is not by itself evidence of prejudice in the item writer; it can arise wherever the same behaviour costs different amounts in different settings. And invariance holding is not a guarantee that a difference is interesting, important, or caused by anything in particular.
The longer version lets you move every loading and intercept separately, walks the full configural, metric and scalar ladder, and shows the item characteristic lines for all four items in both groups. It is at Measurement-Invariance Translator in the main collection.