[DSRP Evidence](https://dsrpevidence.org/)

# Can mental fitness and its mechanisms be measured?

## The research claim

Measuring thinking is where most frameworks stop. A systematic review of the mental fitness literature found six distinct measurement modes with no dominant standard, and a separate review of systems thinking frameworks found proliferation without empirical differentiation. Most accounts of good thinking cannot say how much of it someone has. The claim is that it can be measured, and that the thing being measured is a skill rather than a trait. That distinction decides everything downstream. A psychometric test is built to be stable — improvement would be error. An edumetric test is built to move, because the thing it measures is supposed to change. What is measured directly is the cognitive term: how well a person organizes information, and how accurately they judge their own organizing. The other three domains of fitness are specified as a measurement interface rather than a battery, with criteria for what counts as an admissible indicator. Three properties have to hold. The measure must be stable across occasions. It must relate to what it claims to measure rather than to something already measured by other means. And it must move when the person does — which for an edumetric instrument is the whole point, and the property with the least evidence behind it.

## What would confirm it

Mental fitness and the moves that produce it can be measured — stably, in a way that relates to what is being claimed, and sensitively enough to register change. Read in symbols: reliability meets a stated threshold (τᵣ), validity meets a stated threshold (τᵥ), and sensitivity to change is greater than zero. All three are joined by ∧ because all three are required. An instrument that is stable and unrelated to anything is useless; so is one that relates to everything and never moves.

## What would refute it

The measure is unstable, or measures something other than what it claims, or does not move when the person improves. Read in symbols: reliability below threshold, or (∨) validity below threshold, or sensitivity to change equal to zero. Any one of the three is fatal, which is why they are joined by or. The third is the most commonly fatal in practice: instruments that are stable and valid and register nothing when someone actually learns.

## The evidence, and why this status

Two instruments, ten years apart, and the honest summary is that structure and reliability hold while sensitivity and independence do not. The current instrument shows high internal consistency and excellent model fit, with most reliable variance loading on a general factor and smaller pattern-specific subskills alongside it. That shape is itself a theoretical result rather than a psychometric convenience: the prediction was that the four patterns operate together rather than separately, so one dominant factor with real but secondary specialization is what should appear. A different shape would have been a problem for the theory, not just for the instrument. The weak points are reported rather than buried. Subscale reliabilities in the earlier and larger validation are modest, one pattern has never scored well in any version, and one fit index sits above its conventional threshold. The papers state plainly that the instrument has not been validated against IQ, aptitude tests, or existing metacognition and critical thinking measures — so incremental validity is claimed nowhere. One convergence is worth separating out because neither study was designed to produce it. The instrument finds confidence exceeding skill in every domain, by the widest margin for perspective. Separately, and by a completely different method, perspective is among the least-used patterns when people act with no instruction at all. People are worst at the move they are most confident about, found twice, by designs that fail in different ways. Underneath sits the item-level work: eighteen experiments establishing that each element can be elicited and scored on its own, which is what the instrument counts.

## Supporting evidence (11 publications)

- [Cabrera D, Cabrera L (2025). The Thinking Quotient (TQ): Updated Psychometric Validation and Domain Reliability (2025 dataset). Journal of Systems Thinking (in press).](https://dsrpevidence.org/paper/505)

- [Cabrera L, Sokolow J, Cabrera D (2023). Developing and Validating a Measurement of Systems Thinking (STMI). Journal of Systems Thinking 3(1):1-43.](https://dsrpevidence.org/paper/373)

- [Cabrera D, Cabrera L (2025). The Pareto Structure of Thought: empirical discovery of the six foundational mental moves. Journal of Systems Thinking.](https://dsrpevidence.org/paper/500)

- [Cabrera D, Cabrera L, Cabrera E (2022). Distinctions Organize Information in Mind and Nature (D). Systems 10(2):41.](https://dsrpevidence.org/paper/318)

- [Cabrera D, Cabrera L, Cabrera E (2022). Systems Organize Information in Mind and Nature (S). Systems 10(2):44.](https://dsrpevidence.org/paper/345)

- [Cabrera D, Cabrera L, Cabrera E (2022). Relationships Organize Information in Mind and Nature (R). Systems 10(3):71.](https://dsrpevidence.org/paper/338)

- [Cabrera D, Cabrera L, Cabrera E (2022). Perspectives Organize Information in Mind and Nature (P). Systems 10(3):52.](https://dsrpevidence.org/paper/337)

- [Cabrera, D., Lakshmanaprasath, S., Silberman, D., & Cabrera, L. (2025). Mental fitness and the structure of adaptation: Evidence from a systematic review toward a measurable, trainable science of mental fitness. Journal of Systems Thinking,.](https://dsrpevidence.org/paper/478)

- [Cabrera, D., & Cabrera, L. (2026). From zoo smart to jungle-smart: Training mental fitness for adapting to real-world complexity. Journal of Systems Thinking.](https://dsrpevidence.org/paper/572)

- [Cabrera D, Cabrera L (2026). Measuring Context as Structural Organization. Journal of Systems Thinking 6(1).](https://dsrpevidence.org/paper/604)

- [Cabrera D (2025). A General, Measurable Definition of Emergence (E = ΔO). Journal of Systems Thinking 5(1):1-23.](https://dsrpevidence.org/paper/457)

[All research questions](https://dsrpevidence.org/claims)
