Personality and Individual Differences

Intelligence-Test Battery Builder

Seven task types, one session, three different purposes. Build a battery and watch every choice cost you something else.

Simulated — generic task families, illustrative values, no real test

Learning objective

By the end you should be able to name the trade-offs that assembling a battery forces, and why there is no best battery in the abstract.

About 25 minutes. Nothing you do here is saved or sent anywhere.

What this is not

This tool scores batteries, never people. It produces no ability score for anybody, reproduces no published test, item, norm or scoring rule, and cannot be used to assess or screen anyone. Real cognitive assessment is carried out by qualified practitioners using standardised instruments under conditions none of which are simulated here.

First, a judgement

What makes one test battery better than another?

The builder

Key terms
Purpose
What the assessment is for. Validity attaches to a use, not to a task considered on its own.
Reliability
How consistently a task measures whatever it measures. It is a precondition rather than an achievement.
Coverage
How much of what you are trying to assess the battery reaches between them.
Ceiling
The point above which a task can no longer tell people apart, because they are all near the top of it.
Norms
The reference sample a score is expressed against. A score has no meaning without one, and the choice of sample is substantive.

Assemble a battery

Seven generic task families. Every value shown is illustrative and was written for this tool; none is a measured property of any published instrument.

Requirements

Purpose

Available tasks

    What you have built

    Ability domains sampled
    How much of each ability domain the selected battery samples
    Domain Coverage Status

    What this battery is trading away

    What this demonstrates

    Validity is a property of a use, not of a test

    The same battery scores differently against the three purposes without a single task changing. This is the modern conception of validity: a test is not valid in general, it is valid for a stated inference in a stated context. "Is this a good test?" has no answer until someone says what it is for.

    Every constraint is real and they conflict

    Time, breadth, reliability, exposure-dependence and burden cannot be maximised together. Adding tasks buys coverage and reliability and spends time and burden. Choosing abstract tasks lowers exposure-dependence and usually costs you the verbal and quantitative domains entirely. There is no configuration that wins on everything, which is why real test construction is a series of defensible compromises rather than an optimisation.

    Reliability is the cheapest thing to buy and the least informative

    Composite reliability rises with the number of tasks almost regardless of what they measure. It is entirely possible to build a battery with excellent reliability that samples two of six domains — and that combination reads as rigorous in a methods section. Coverage is harder to buy and matters more.

    Removing language reduces some demands and removes no culture

    Matrix and rotation tasks carry lower exposure-dependence than vocabulary, and not zero. Every task here still requires familiarity with being tested, with the response format, with working quickly because speed is wanted, and with the convention that abstract puzzles have single correct answers. "Culture-fair" names an aspiration and a direction of travel, not an achieved property.

    Download activity HTML

    For teaching elsewhere: take this activity as one self-contained block of HTML, on the clipboard or as a file. Either way it is styled so that it will not disturb the page you put it into.