Personality & Individual Differences · Simplified
Take the Words Out and See What Is Left
A test with no words in it is not a test with no assumptions in it.
How this handles groups
How this handles groups. What is modelled here is opportunity to learn a test format. The three fictional people are described only by their prior experience of testing formats, devices and puzzle conventions. None has a nationality, ethnicity, region, language or group membership, no score is produced for anyone, and no difference between any groups is simulated. Attaching unfamiliarity with a format to real populations would teach something else, and something false.
Step 1 of 2
First, a judgement
A test designer takes a reasoning test and removes every word from it. No written instructions, no verbal items, nothing to read. What has that achieved?
Step 2 of 2
Four decisions
You are designing a reasoning task. Every decision below changes what the task demands of somebody beyond the reasoning it is meant to measure. The design you start with is a poor one.
Key terms
- Construct-relevant demand
- Something the task requires that is part of what you mean to measure. Here, reasoning.
- Construct-irrelevant demand
- Something the task requires that is not. Differences in score produced by it are differences in something you did not intend to measure.
- Opportunity to learn
- Prior exposure to the format and conventions a task uses, as distinct from the ability it is meant to tap.
- Comparability
- Evidence that scores mean the same thing in the groups being compared. It is established rather than assumed.
One bar per demand that is not reasoning. A short bar means the task asks little of that kind. The marked line is where a demand stops being negligible.
| Prior experience | Demand falling on them | Where most of it comes from |
|---|
What this shows
Fairness is a direction, not a destination
Key idea: Taking the words out did what it says on the tin. The reading demand went from the top of the chart to the bottom, and that is a real improvement worth making. Two other demands went up while it happened, because abstract figure puzzles are themselves a format somebody has either met before or has not, and working out when to guess and when to move on is a skill picked up by sitting tests. The best design available here still leaves every demand above zero, and the one that never falls far is the most ordinary of all: knowing how to be a person sitting a test. That is why culture-fair is better described as a direction than as a property a test can have. The practical conclusion is not that testing is hopeless. It is that reducing an irrelevant demand and showing that scores are comparable are two different jobs, and doing the first does not do the second.
The demand levels are invented numbers chosen to make the trade-offs legible, not measurements of anything, and the six demands are not a complete list. Three cautions about the argument. Showing that no design removes every irrelevant demand is not an argument against testing, and the differences between the designs here are large: the worst is far worse than the best, and choosing well matters. What falls on the three fictional people is a property of the design meeting a history of exposure, and it is not a score, a prediction or an estimate of anybody's reasoning. And nothing here models any real group: the whole point of describing people only by what formats they have met is that unfamiliarity with a format is not a property of a population.
The longer version offers six decisions rather than four, a fourth fictional participant, an explicit construct-relevant share of the score, and a closing exercise on what evidence comparability would require. It is at Culture-Fair Test Challenge in the main collection.