Vantry is an invented country with two long-established regional populations: Uplanders, about 18 per cent of the population, and Lowlanders, about 82 per cent. Uplanders hold fewer senior posts than their share of the workforce would suggest. A research team has five instruments and a question, and the question is not "how much prejudice is there" - it is what would each of these actually be evidence of.
Two hundred simulated Lowlander respondents did two things. They sorted words and images under two pairings, producing a latency difference score in seconds - larger means faster under one pairing than the other. Then they each chose four candidates from eight for a fictional post, four of whom were coded as Uplanders. Three of the two hundred are picked out below. You see each score first, and the scatter beside you always shows all 200.
Nothing here changes the controls for you, and none of it says what you are going to find.
Three respondents, one at a time. Nothing is timed, and there is no penalty for saying you cannot tell - that answer is treated seriously.
| Uplanders shortlisted | Respondents | Mean latency score | Range |
|---|
Five findings from Vantry, each produced by a different method. For each one, the question is not whether it shows prejudice. It is what the instrument put in front of you, before anybody interpreted it.
Each answer is followed by what is being inferred, what remains unexplained, and which level of description the finding belongs to.
Progress
| Instrument | Directly observes | Level | Cannot support |
|---|
Five conclusions somebody might draw from the Vantry evidence. Say whether each is supported, partly supported, or not supported.
The correlation between the latency score and the shortlist in this simulated sample is small, real, and useless about any particular person. Both things are true at once, and holding them together is the whole skill. A small association tells you something about how a population is arranged; it does not tell you which side of the line the person in front of you is on. Using the score to say anything about an individual - to select, to train, to reassure, to accuse - is a category error, and no amount of statistical significance repairs it.
It is a difference in mean sorting speed between two pairings of categories. That is a real observation about how quickly certain associations become available under time pressure, and it is interesting. It is not a reading taken from somewhere behind the person's answers. Familiarity with the stimuli, the order the blocks came in, how salient each category is and general processing speed all contribute to the number, and none of them is an attitude. Calling it "the real attitude, underneath" is a rhetorical move rather than a measurement claim.
What somebody reports on a form, what their sorting speed does, what they did on one occasion, how they appear to two colleagues, and what an organisation's appointment record looks like over twelve years are five findings at four levels of description. Each is evidence of itself. The traffic between them is where the trouble is: an appointments gap does not become explained because a sample of individuals showed an association effect, and a population of individuals with no measurable association effect would not thereby be shown to have a fair appointments process.
A survey answer is often dismissed as contaminated by social desirability, and then a latency score is treated as the honest version. But what people are willing to say is itself worth knowing, and a narrowing self-report gap over ten years might mean attitudes changed, or that what is sayable changed, or both - and that second possibility is a finding about a society, not noise in a measurement. The instrument observes what people will write down under these conditions. That is a real object of study.
Every one of the five findings is a trace left by a procedure. The moment a research programme starts saying "prejudice" when it means "the score our task produces", the construct has quietly been replaced by the instrument, and every subsequent finding is a finding about the instrument. That is not a hypothetical failure mode. It is the ordinary drift of a successful measure, and the defence against it is exactly the question this tool keeps asking: what did the procedure put in front of us, before anybody interpreted it?