Personality and Individual Differences
Culture-Fair Test Challenge
Design a reasoning task with no words in it, and watch how much of what it asks for still has nothing to do with reasoning.
Simulated — a design model, not a test, and not about any real group
Learning objective
By the end you should be able to tell construct-relevant demands from construct-irrelevant ones, and say what is required before scores can be compared across groups.
About 25 minutes. Nothing you do here is saved or sent anywhere.
How this handles groups
This tool models opportunity to learn a test format. The four fictional participants are described only by their prior experience of testing formats, devices and puzzle conventions. None has a nationality, ethnicity, region, language or group membership, and the tool never produces a score for anyone or simulates a difference between any groups. Attaching format unfamiliarity to real populations would teach something else, and something false.
First, a judgement
A test designer takes a reasoning test and removes every word from it — no written instructions, no verbal items, nothing to read.
The designer
Six decisions. The design you start with is a poor one — change it and watch the demands move.
Key terms
- Construct-relevant demand
- Something the task requires that is part of what you mean to measure.
- Construct-irrelevant variance
- Differences in score produced by demands that are not part of the construct.
- Opportunity to learn
- Prior exposure to the format and conventions a task uses, as distinct from the ability it is meant to tap.
- Comparability
- Evidence that scores mean the same thing in the groups being compared. It is established rather than assumed.
A non-verbal reasoning task
Every figure below is model-implied from illustrative values written for this tool. No published test is described, and no score is produced for anybody.
What your task demands
| Demand | Level | Rating |
|---|
What the score is made of
Four fictional people take your task
Each is described only by their prior experience of testing formats, devices and puzzle conventions. The figure shown is the construct-irrelevant load your design places on them: how much of what your task asks for is something they have not had the chance to become familiar with.
It is not a score, and not an estimate of how well they reason. Nothing here predicts anybody's performance.
Challenge — what would you need to know?
Someone wants to compare average scores on your task between two groups of people who were educated in different systems. Which of these would count as evidence that the scores are comparable? Select all that apply.
What this demonstrates
Removing language removes a demand, not the problem
Reading is one of six demands modelled here, and the only one that disappears when the words go. Familiarity with formal testing as a genre, with abstract puzzles as a kind of object, with the expectation that speed is wanted, with the response medium, and with when to guess — all survive intact. "Non-verbal" is a description of the items, not a claim about the demands.
Construct-irrelevant variance is the useful concept
The question is not whether a test is "fair" in the abstract but how much of its score reflects the thing it is meant to measure and how much reflects something else. That framing is testable, improvable and specific, where "culture-free" is none of those.
Opportunity to learn is about the format, not the person
Everything modelled here concerns whether someone has previously encountered a way of being tested. That is a fact about what somebody has been exposed to. It says nothing about capacity, and treating it as though it did is the inferential error to avoid.
Comparability is established, not assumed
Before two groups' scores can be compared, someone has to show that the test measures the same thing in both — through measurement invariance testing, differential item functioning analysis, evidence about prior familiarity, and qualitative work on how people actually approached the items. Until that exists, the honest position is that the scores may not be on the same scale. That is a statement about the instrument.
Fairness costs something
The lowest-demand design is untimed, demonstrated rather than explained, and preceded by practice with feedback. It takes far longer to administer. Every real testing programme trades this against cost, and the trade is a decision someone makes rather than a constraint they discover.
For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.