The setting
Vantry is an invented country with two long-established regional populations: Uplanders, about 18 per cent of the population, and Lowlanders, about 82 per cent. Uplanders hold fewer senior posts than their share of the workforce would suggest. A research team has five instruments and a question, and the question is not "how much prejudice is there" - it is what would each of these actually be evidence of.
Experiment 1 - a score, and a person
Two hundred simulated Lowlander respondents did two things. They sorted words and images under two pairings, producing a latency difference score in seconds - larger means faster under one pairing than the other. Then they each chose four candidates from eight for a fictional post, four of whom were coded as Uplanders. Three of the two hundred are picked out below. You see each score first, and the scatter beside you always shows all 200.
Need a hand?
Nothing here changes the controls for you, and none of it says what you are going to find.
Key terms
- Response latency
- How long somebody takes to respond. A difference between conditions is directly a difference in times.
- Level of description
- Whether a finding is about an individual, an interaction, an institution or a population. Evidence at one level does not carry to another for free.
- Construct
- What you mean to be studying, as distinct from the observable thing standing in for it.
- Average relationship
- An association holding across a sample. It is not a prediction about any member of it.
Predict one person from one score
Three respondents, one at a time. Nothing is timed, and there is no penalty for saying you cannot tell - that answer is treated seriously.
All 200, score against shortlist
The same data as a table
| Uplanders shortlisted | Respondents | Mean latency score | Range |
|---|
Experiment 2 - five instruments
Five findings from Vantry, each produced by a different method. For each one, the question is not whether it shows prejudice. It is what the instrument put in front of you, before anybody interpreted it.
What does this one directly observe?
Each answer is followed by what is being inferred, what remains unexplained, and which level of description the finding belongs to.
Progress
The finding
Five instruments, four levels
| Instrument | Directly observes | Level | Cannot support |
|---|
Challenge - what does the evidence license?
Five conclusions somebody might draw from the Vantry evidence. Say whether each is supported, partly supported, or not supported.
What this demonstrates
An average relationship is not an individual prediction
The correlation between the latency score and the shortlist in this simulated sample is small, real, and useless about any particular person. Both things are true at once, and holding them together is the whole skill. A small association tells you something about how a population is arranged; it does not tell you which side of the line the person in front of you is on. Using the score to say anything about an individual - to select, to train, to reassure, to accuse - is a category error, and no amount of statistical significance repairs it.
What a latency difference actually is
It is a difference in mean sorting speed between two pairings of categories. That is a real observation about how quickly certain associations become available under time pressure, and it is interesting. It is not a reading taken from somewhere behind the person's answers. Familiarity with the stimuli, the order the blocks came in, how salient each category is and general processing speed all contribute to the number, and none of them is an attitude. Calling it "the real attitude, underneath" is a rhetorical move rather than a measurement claim.
Four levels, and the traffic between them
What somebody reports on a form, what their sorting speed does, what they did on one occasion, how they appear to two colleagues, and what an organisation's appointment record looks like over twelve years are five findings at four levels of description. Each is evidence of itself. The traffic between them is where the trouble is: an appointments gap does not become explained because a sample of individuals showed an association effect, and a population of individuals with no measurable association effect would not thereby be shown to have a fair appointments process.
Why self-report is not simply the weaker instrument
A survey answer is often dismissed as contaminated by social desirability, and then a latency score is treated as the honest version. But what people are willing to say is itself worth knowing, and a narrowing self-report gap over ten years might mean attitudes changed, or that what is sayable changed, or both - and that second possibility is a finding about a society, not noise in a measurement. The instrument observes what people will write down under these conditions. That is a real object of study.
The instrument is not the construct
Every one of the five findings is a trace left by a procedure. The moment a research programme starts saying "prejudice" when it means "the score our task produces", the construct has quietly been replaced by the instrument, and every subsequent finding is a finding about the instrument. That is not a hypothetical failure mode. It is the ordinary drift of a successful measure, and the defence against it is exactly the question this tool keeps asking: what did the procedure put in front of us, before anybody interpreted it?
For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.