Personality & Individual Differences · Simplified
No Best Battery
Sixty minutes of testing time, seven things you could do with it, and three different reasons somebody might be sitting there.
What this is not
What this is not. This page scores batteries and never people. It produces no ability score for anybody, reproduces no published test, item, norm or scoring rule, and cannot be used to assess or screen anyone. The task families are generic descriptions, and every number attached to them was written for teaching. Real cognitive assessment is carried out by qualified practitioners using standardised instruments under conditions none of which are simulated here.
Step 1 of 2
First, a judgement
What makes one test battery better than another?
Step 2 of 2
Sixty minutes to spend
Choose the tasks, then choose what the session is for. The two choices are independent, which is the thing to notice: changing the purpose changes the verdict without changing a single task.
Key terms
- Purpose
- What the assessment is for. Validity attaches to a use, not to a task considered on its own.
- Coverage
- How much of what you are trying to reach the battery reaches between its tasks. A battery that measures one thing four ways is not broad, however long it is.
- Composite reliability
- How consistently the total score measures whatever it measures. It is a precondition rather than an achievement.
- Dependence on prior opportunity
- How much a task's score turns on having had the chance to learn its particular content, in a particular language and a particular schooling. It is a property of a task, never of a person or a group.
- Burden
- What the session costs the person sitting it: fatigue, stress, and what it demands of movement, hearing and sight.
One bar per broad ability area. A bar reaching the marked line means the battery samples that area at all. Two tasks measuring the same area do not make the battery broader.
What this shows
The battery did not change. The question did.
Key idea: Every property on this page is a real property of the battery, and not one of them is a verdict. Coverage, composite reliability, dependence on prior opportunity and burden are the same numbers whatever the session is for, and they turn into good or bad only once somebody says what the assessment is being used to decide. That is what it means to say validity belongs to a use rather than to a test. The constraint underneath it is time. Sixty minutes buys you either depth in a few areas or a thin sweep across several, and there is no arrangement of these seven tasks inside the budget that satisfies all three purposes at once. Not one. That is not a failure of the task list; it is what a trade-off is. The practical version of this is that the first question to ask about any assessment is not how good it is but what it is for, and that a battery assembled without answering it will turn out to have answered it by accident.
The task families are generic and every number attached to them was invented for teaching, so the specific trade-offs here are not the trade-offs any real battery faces. Three cautions about the ideas rather than the numbers. Dependence on prior opportunity is a property of a task and never of a person or a group; a task that assumes particular schooling tells you about the task's demands, and reading it as a fact about whoever scores lower on it is exactly the inference this page is built to block. A score also means nothing without the reference sample it is expressed against, and choosing that sample is a substantive decision this page does not model at all. And none of the properties here, alone or together, licenses a claim about anybody's ability: a battery can be well built for its purpose and still be the wrong thing to have done to somebody.
The longer version adds further task families, a burden tolerance that varies by scenario, a graded fit score rather than a checklist, and an explicit treatment of norms and ceilings. It is at Intelligence-Test Battery Builder in the main collection.