Personality and Individual Differences

Reverse-Item Disaster

Six items, two of them reverse-keyed, and four ways of scoring them — one right and three wrong in ways that leave almost no trace.

Simulated — a fictional scale and six fictional respondents

Learning objective

By the end you should be able to recode a reverse-keyed item correctly, and recognise in the item statistics when somebody has not.

About 20 minutes. Nothing you do here is saved or sent anywhere.

The scale

Six original items about persistence, answered 1 to 5. Two are worded so that agreement means less of the trait.

    First, a prediction

    A researcher scores this scale and forgets to recode the two reverse-keyed items.

    What is the most serious consequence?

    What this demonstrates

    The damage is to people, not to the scale

    Forgetting to recode lowers alpha and produces negative item-total correlations, and those are the symptoms usually taught. The consequence that matters is that individual scores are wrong and the respondents come out in the wrong order. Every correlation the scale then enters into is computed on those scores.

    Some errors announce themselves and some do not

    No recoding at all is loud: the reverse items correlate negatively with the rest of the scale, which anybody inspecting item statistics will see. Using the wrong scale maximum is silent. Every reverse answer comes out exactly one point too high, alpha hardly moves, no item looks anomalous, and the error survives peer review comfortably.

    Reading an item as reversed is not recoding it

    A researcher who understands perfectly well that "I give up on things once they become difficult" is reverse-worded, and leaves the number alone, has made exactly the same arithmetic error as somebody who never noticed. Comprehension and recoding are different operations, and only one of them is in the data.

    Reverse wording is a trade, not a free fix

    Reverse items exist to counter acquiescence, and they work — the Response-Style Simulator in this module shows how. They also cost something. They are harder to understand, they are answered inconsistently by people reading quickly, they attract careless responding, and they routinely form their own factor in a factor analysis, which then gets mistaken for a second substantive dimension. Including them is a judgement, not an obligation.