Personality and Individual Differences
Design a reasoning task with no words in it, and watch how much of what it asks for still has nothing to do with reasoning.
Simulated — a design model, not a test, and not about any real group
By the end you should be able to tell construct-relevant demands from construct-irrelevant ones, and say what is required before scores can be compared across groups.
About 25 minutes. Nothing you do here is saved or sent anywhere.
This tool models opportunity to learn a test format. The four fictional participants are described only by their prior experience of testing formats, devices and puzzle conventions. None has a nationality, ethnicity, region, language or group membership, and the tool never produces a score for anyone or simulates a difference between any groups. Attaching format unfamiliarity to real populations would teach something else, and something false.
A test designer takes a reasoning test and removes every word from it — no written instructions, no verbal items, nothing to read.
Six decisions. The design you start with is a poor one — change it and watch the demands move.
Every figure below is model-implied from illustrative values written for this tool. No published test is described, and no score is produced for anybody.
| Demand | Level | Rating |
|---|
Each is described only by their prior experience of testing formats, devices and puzzle conventions. The figure shown is the construct-irrelevant load your design places on them: how much of what your task asks for is something they have not had the chance to become familiar with.
It is not a score, and not an estimate of how well they reason. Nothing here predicts anybody's performance.
Someone wants to compare average scores on your task between two groups of people who were educated in different systems. Which of these would count as evidence that the scores are comparable? Select all that apply.
Reading is one of six demands modelled here, and the only one that disappears when the words go. Familiarity with formal testing as a genre, with abstract puzzles as a kind of object, with the expectation that speed is wanted, with the response medium, and with when to guess — all survive intact. "Non-verbal" is a description of the items, not a claim about the demands.
The question is not whether a test is "fair" in the abstract but how much of its score reflects the thing it is meant to measure and how much reflects something else. That framing is testable, improvable and specific, where "culture-free" is none of those.
Everything modelled here concerns whether someone has previously encountered a way of being tested. That is a fact about what somebody has been exposed to. It says nothing about capacity, and treating it as though it did is the inferential error to avoid.
Before two groups' scores can be compared, someone has to show that the test measures the same thing in both — through measurement invariance testing, differential item functioning analysis, evidence about prior familiarity, and qualitative work on how people actually approached the items. Until that exists, the honest position is that the scores may not be on the same scale. That is a statement about the instrument.
The lowest-demand design is untimed, demonstrated rather than explained, and preceded by practice with feedback. It takes far longer to administer. Every real testing programme trades this against cost, and the trade is a decision someone makes rather than a constraint they discover.