Personality and Individual Differences
Intelligence-Test Battery Builder
Seven task types, one session, three different purposes. Build a battery and watch every choice cost you something else.
Simulated — generic task families, illustrative values, no real test
Learning objective
By the end you should be able to name the trade-offs that assembling a battery forces, and why there is no best battery in the abstract.
About 25 minutes. Nothing you do here is saved or sent anywhere.
What this is not
This tool scores batteries, never people. It produces no ability score for anybody, reproduces no published test, item, norm or scoring rule, and cannot be used to assess or screen anyone. Real cognitive assessment is carried out by qualified practitioners using standardised instruments under conditions none of which are simulated here.
First, a judgement
The builder
Key terms
- Purpose
- What the assessment is for. Validity attaches to a use, not to a task considered on its own.
- Reliability
- How consistently a task measures whatever it measures. It is a precondition rather than an achievement.
- Coverage
- How much of what you are trying to assess the battery reaches between them.
- Ceiling
- The point above which a task can no longer tell people apart, because they are all near the top of it.
- Norms
- The reference sample a score is expressed against. A score has no meaning without one, and the choice of sample is substantive.
Assemble a battery
Seven generic task families. Every value shown is illustrative and was written for this tool; none is a measured property of any published instrument.
Requirements
What you have built
| Domain | Coverage | Status |
|---|
What this battery is trading away
Your battery, judged three ways
| Purpose | Fit | Session time | Burden |
|---|
What this demonstrates
Validity is a property of a use, not of a test
The same battery scores differently against the three purposes without a single task changing. This is the modern conception of validity: a test is not valid in general, it is valid for a stated inference in a stated context. "Is this a good test?" has no answer until someone says what it is for.
Every constraint is real and they conflict
Time, breadth, reliability, exposure-dependence and burden cannot be maximised together. Adding tasks buys coverage and reliability and spends time and burden. Choosing abstract tasks lowers exposure-dependence and usually costs you the verbal and quantitative domains entirely. There is no configuration that wins on everything, which is why real test construction is a series of defensible compromises rather than an optimisation.
Reliability is the cheapest thing to buy and the least informative
Composite reliability rises with the number of tasks almost regardless of what they measure. It is entirely possible to build a battery with excellent reliability that samples two of six domains — and that combination reads as rigorous in a methods section. Coverage is harder to buy and matters more.
Removing language reduces some demands and removes no culture
Matrix and rotation tasks carry lower exposure-dependence than vocabulary, and not zero. Every task here still requires familiarity with being tested, with the response format, with working quickly because speed is wanted, and with the convention that abstract puzzles have single correct answers. "Culture-fair" names an aspiration and a direction of travel, not an achieved property.
For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.