Personality and Individual Differences
Start with a built five-item scale, add one paraphrase at a time, and watch Cronbach’s alpha climb while the scale quietly stops measuring most of what it was for.
Simulated data — fictional items, not a real questionnaire
By the end you should be able to say what a high alpha does and does not tell you about a scale.
About 15 minutes. Nothing you do here is saved or sent anywhere.
You are about to be given a built five-item scale measuring “academic conscientiousness”, and the only thing you can do to it is add another way of asking a question it already contains. Commit to an answer now — the point of the exercise is largely lost if you read the outcome first.
The scale is already built. Every figure below is calculated from a fictional correlation model, not from real respondents.
Now Record your prediction above to open the scale.
| Content area | Items | Repeated wording |
|---|
No items selected yet.
A fictional outcome — tutor-rated dependability — generated by the same model. It is an illustration of a principle, not a validity coefficient from any real study.
| Measure | With every repeat | Repeats dropped |
|---|
The first can be attempted at any point. The second uses your repaired scale, so it unlocks once the additions are done.
Cronbach’s alpha rises with the number of items and with the average correlation between them. Neither is a fact about the construct. Write an item five times over and both go up at once. That is how a scale of paraphrases reaches 0.91 while covering a fifth of the ground the construct occupies.
A high alpha is often read as “the items measure one thing”. It does not establish that. Alpha can be high for a clearly multidimensional scale, simply because it has many items; and it can be high because one narrow cluster dominates. Deciding how many dimensions there are is a question for factor analysis, not for alpha.
Reliability limits validity but does not supply it. In this simulation, adding near-duplicates loads the total score with variance specific to one phrasing — variance the outcome does not care about. The scale becomes a more precise measure of something narrower. That is the bloated specific factor, and it is a real failure mode in published scales, not an artefact of this demonstration.
For a broad construct, an alpha above about 0.90 in a short scale usually means the items repeat each other. The useful question is never “is alpha high enough?” but “is this scale made of items that between them cover what I claim to be measuring?”