Personality and Individual Differences

The Alpha Trap

Start with a built five-item scale, add one paraphrase at a time, and watch Cronbach’s alpha climb while the scale quietly stops measuring most of what it was for.

Simulated data — fictional items, not a real questionnaire

Learning objective

By the end you should be able to say what a high alpha does and does not tell you about a scale.

About 15 minutes. Nothing you do here is saved or sent anywhere.

How to use this

  1. Commit to a prediction before you open the scale.
  2. Add another way of asking a question the scale already contains. Watch all three numbers, not just alpha.
  3. Do it four more times, then read what the scale now measures.
  4. Drop the repeats, and compare the two scales.
  5. Work through the two challenges at the foot of the page.

First, a prediction

You are about to be given a built five-item scale measuring “academic conscientiousness”, and the only thing you can do to it is add another way of asking a question it already contains. Commit to an answer now — the point of the exercise is largely lost if you read the outcome first.

1. The construct has five content areas. How many do you think your alpha-chasing scale will end up covering?
2. Suppose we then correlate that scale with a relevant outcome. What would you expect?

The scale builder

Key terms
Cronbach's alpha
An index of how consistently a set of items covary. It is a function of two things: how many items there are, and how strongly they correlate on average.
Mean inter-item correlation
How strongly the items agree with each other, averaged over every pair.
Unidimensionality
Whether the items are measuring one thing. It is a different question from internal consistency.
Construct validity
Whether the scale measures what it claims to. Answering it needs evidence from outside the scale itself.
Near-duplicate items
Items whose wording is almost the same as each other's.

Academic conscientiousness — a five-item scale

The scale is already built. Every figure below is calculated from a fictional correlation model, not from real respondents.

Now Record your prediction above to open the scale.

The items in your scale
    Change the scale

    Your scale

    Cronbach’s alpha
    Items
    0
    Mean inter-item r
    Content areas covered
    0 of 5
    Effective content areas
    Near-duplicate pairs
    0
    Content coverage
    Items selected in each content area, and how many of those repeat the wording of another selected item
    Content area Items Repeated wording

    Redundancy

    No items selected yet.

    Two challenges

    The first can be attempted at any point. The second uses your repaired scale, so it unlocks once the additions are done.

    Challenge 1 — diagnose three measures

    Three fictional research groups publish a scale for the same construct. Which is the strongest measure of it?

    Three fictional scales compared
    Scale Items Alpha Content areas Largest set of near-identical items
    A 120.942 of 57 items
    B 100.795 of 51 item
    C 60.515 of 51 item
    Your answer

    Challenge 2 — predict the effect

    This unlocks once you have made all five additions. It asks you to predict what happens to a repaired scale when four near-duplicate items are added back to it.

    What this demonstrates

    Alpha is a function of two things you control

    Cronbach’s alpha rises with the number of items and with the average correlation between them. Neither is a fact about the construct. Write an item five times over and both go up at once. That is how a scale of paraphrases reaches 0.91 while covering a fifth of the ground the construct occupies.

    Internal consistency is not unidimensionality

    A high alpha is often read as “the items measure one thing”. It does not establish that. Alpha can be high for a clearly multidimensional scale, simply because it has many items; and it can be high because one narrow cluster dominates. Deciding how many dimensions there are is a question for factor analysis, not for alpha.

    Internal consistency is not validity

    Reliability limits validity but does not supply it. In this simulation, adding near-duplicates loads the total score with variance specific to one phrasing — variance the outcome does not care about. The scale becomes a more precise measure of something narrower. That is the bloated specific factor, and it is a real failure mode in published scales, not an artefact of this demonstration.

    Very high alpha is a warning, not a prize

    For a broad construct, an alpha above about 0.90 in a short scale usually means the items repeat each other. The useful question is never “is alpha high enough?” but “is this scale made of items that between them cover what I claim to be measuring?”

    For educators

    The teaching notes record the intended level, a seminar sequence, debrief questions, likely misconceptions, and the exact simulation model.

    Read the teaching notes

    Download activity HTML

    For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.