Social and Critical Psychology

Measuring Prejudice: What Does the Instrument Capture?

A survey answer, a difference in sorting speed, a shortlist, a colleague's impression and twelve years of appointment figures. Five instruments, all pointed at the same word, all observing something different from it - and sitting at four different levels that cannot be run together.

Entirely simulated - Vantry, its two regions and all 200 respondents are generated by a documented model

Learning objective

By the end you should be able to say what each instrument directly observes, and where an instrument stops being evidence about a construct and becomes the construct.

About 20 minutes. Nothing is timed, and nothing you select is saved or sent anywhere.

Nothing here tests you

There is no task here that measures anything about the person using it. The sorting task is run on 200 fictional respondents generated by a seeded model, and you never take it. Vantry, the Uplands and the Lowlands are invented, no group in them stands for any real group anywhere, and no figure is a prevalence estimate, a norm or a published effect.

How to use this

  1. Read the setting and commit to an answer to the question below.
  2. Experiment 1: for each of three fictional respondents, read the sorting score, predict what they did in a shortlisting task, and find out.
  3. Experiment 2 opens after the third. Take each of the five instruments in turn and say what it directly observes.
  4. Then judge five conclusions somebody might draw.

The setting

Vantry is an invented country with two long-established regional populations: Uplanders, about 18 per cent of the population, and Lowlanders, about 82 per cent. Uplanders hold fewer senior posts than their share of the workforce would suggest. A research team has five instruments and a question, and the question is not "how much prejudice is there" - it is what would each of these actually be evidence of.

Two hundred people take a latency-based sorting task and then shortlist candidates for a fictional post. How well will the sorting score predict any individual's shortlist?

Challenge - what does the evidence license?

Five conclusions somebody might draw from the Vantry evidence. Say whether each is supported, partly supported, or not supported.

What this demonstrates

An average relationship is not an individual prediction

The correlation between the latency score and the shortlist in this simulated sample is small, real, and useless about any particular person. Both things are true at once, and holding them together is the whole skill. A small association tells you something about how a population is arranged; it does not tell you which side of the line the person in front of you is on. Using the score to say anything about an individual - to select, to train, to reassure, to accuse - is a category error, and no amount of statistical significance repairs it.

What a latency difference actually is

It is a difference in mean sorting speed between two pairings of categories. That is a real observation about how quickly certain associations become available under time pressure, and it is interesting. It is not a reading taken from somewhere behind the person's answers. Familiarity with the stimuli, the order the blocks came in, how salient each category is and general processing speed all contribute to the number, and none of them is an attitude. Calling it "the real attitude, underneath" is a rhetorical move rather than a measurement claim.

Four levels, and the traffic between them

What somebody reports on a form, what their sorting speed does, what they did on one occasion, how they appear to two colleagues, and what an organisation's appointment record looks like over twelve years are five findings at four levels of description. Each is evidence of itself. The traffic between them is where the trouble is: an appointments gap does not become explained because a sample of individuals showed an association effect, and a population of individuals with no measurable association effect would not thereby be shown to have a fair appointments process.

Why self-report is not simply the weaker instrument

A survey answer is often dismissed as contaminated by social desirability, and then a latency score is treated as the honest version. But what people are willing to say is itself worth knowing, and a narrowing self-report gap over ten years might mean attitudes changed, or that what is sayable changed, or both - and that second possibility is a finding about a society, not noise in a measurement. The instrument observes what people will write down under these conditions. That is a real object of study.

The instrument is not the construct

Every one of the five findings is a trace left by a procedure. The moment a research programme starts saying "prejudice" when it means "the score our task produces", the construct has quietly been replaced by the instrument, and every subsequent finding is a finding about the instrument. That is not a hypothetical failure mode. It is the ordinary drift of a successful measure, and the defence against it is exactly the question this tool keeps asking: what did the procedure put in front of us, before anybody interpreted it?