Can a Machine Know What It Does Not Know? Metacognition as the Missing Capacity

MindHeaven® Research DeskEdited by Nikos DrosakisPublished
Preliminary evidence
Narrative review and scientific commentary5 min read3 references

Abstract

Our other work in this department asks whether a machine could have experiences. This article covers a different question, which is more tractable and arguably more urgent: whether a machine can know the limits of its own competence.

Eleven researchers — psychologists, computer scientists and machine learning specialists including Yoshua Bengio, Melanie Mitchell and Bernhard Schölkopf — argue that this is the capacity current systems most conspicuously lack, and that supplying it would matter more than further gains in raw capability.

Their framing is unusual for the field. They call the missing thing wisdom, define it precisely enough to argue about, and separate it carefully from anything to do with consciousness.

1.The Distinction Worth Having

Samuel Johnson and Igor Grossmann at Waterloo, with Amir-Hossein Karimi, Bengio, Nick Chater, Tobias Gerstenberg, Kate Larson, Sydney Levine at Google DeepMind, Mitchell at the Santa Fe Institute, Iyad Rahwan and Schölkopf at the Max Planck Institutes, set out the argument in a paper accepted by Trends in Cognitive Sciences.

They begin from intractable problems — those outside the scope of analytic techniques, where no procedure delivers the right answer and judgement is required. Most consequential human problems are of this kind.

For these, they distinguish two levels of strategy. Object-level strategies yield candidate solutions: heuristics, rules of thumb, the things you actually do. Metacognitive strategies manage the fit between an object-level strategy and the task in front of you.

That second level is the paper's subject, and their examples are concrete: intellectual humility, perspective-taking, context-adaptability. Knowing which of your methods applies here, how confident to be, and whose view you are missing.

Their claim is that AI systems particularly struggle with this type of metacognition — not with generating candidate answers, which they do prolifically, but with judging whether the approach fits the problem.

2.Why This Is Not the Consciousness Question

It is worth being explicit, because the two get conflated constantly and the conflation runs in both directions.

Metacognition is a system's capacity to model and regulate its own processing. It is behaviourally measurable: you can test whether a system's confidence tracks its accuracy, whether it declines problems outside its competence, whether it revises when given contrary evidence.

Phenomenal consciousness — whether there is something it is like to be the system — is not measurable that way, which is the whole reason the indicator method we covered separately had to be invented.

A system could have excellent metacognition and no experience whatsoever. A thermostat models its own state in the most rudimentary way and nobody thinks it feels the cold. Conversely, higher-order theories of consciousness hold that representing one's own mental states is part of what makes them conscious — so the two topics touch, without one settling the other.

Johnson and colleagues stay firmly on the measurable side. Nothing in this paper claims machines have or could have experiences, and it should not be cited as though it did.

3.What Better Metacognition Would Buy

They list four consequences, and the list is notable for being unglamorous.

Robustness to novel environments — a system that recognises when a situation is outside what it has seen can behave cautiously rather than confidently wrong. Explainability — a system with a model of its own reasoning has something to report. Cooperation — shared goals require representing what a partner knows and intends. And safety, through fewer misaligned goals, which follows from the other three.

The word doing the most work across all four is confidence. Overconfidence appears repeatedly in their diagnosis of current failures, and it links directly to misinformation propagation: a system that does not know what it does not know will assert either way with the same fluency.

That is a description most people who use these systems will recognise, and it is worth noting that the failure is not a shortage of knowledge. It is the absence of a model of the boundary of the knowledge.

4.What They Propose

Three directions: benchmarking wisdom, training wise reasoning strategies, and adapting architecture for metacognition. The third is the most substantive suggestion — having systems consider several metacognitive queries alongside the primary query, rather than answering the question and stopping.

The benchmarking problem is where they are most candid about difficulty. Existing evaluations do not capture the richness of everyday intractable problems that wise judgment handles — which is close to definitional, because a problem with a scorable answer is not intractable in their sense.

How to benchmark judgement without reducing it to something that has a right answer is not solved here. It is named as the obstacle.

5.What This Paper Is

A programmatic argument, not a result. There are no experiments, no systems built, no measurements. The authors sketch a vision and propose a research direction.

Its authority comes from the composition of the author list rather than from data — which is a real form of evidence about what a field considers worth pursuing, and no evidence at all about whether the approach will work.

The version we read is the accepted manuscript, posted by the authors ahead of publication. We flag that as we would any preprint, though the acceptance means it has been through review.

What we would want before treating any of this as established is a benchmark that survives contact with a system optimised against it — which is precisely the problem the authors identify and do not solve.

Editorial Comment

MindHeaven® sells supplements. We have no AI product and nothing at stake in whether machines acquire good judgement.

We cover this because the distinction at its centre is one we use constantly without having had a name for it. Object-level strategy is deciding what a study shows. Metacognitive strategy is deciding whether the method you are using to read studies fits this study — whether an effect size means what you are taking it to mean, whether a null is informative or merely underpowered.

Every error we have had to correct in this library has been at the second level rather than the first. Not misreading a result, but applying a reading procedure that did not fit — assuming a source was unavailable without testing it, taking a single trial's direction as the literature's direction.

Intellectual humility, in this paper's sense, is not modesty. It is an accurate model of where your method stops working. That is a demanding standard and we do not consistently meet it.

How to read this article
Preliminary evidence

Mechanism or early findings only — largely animal, cell or unpublished work.

  1. 1.Johnson SGB, Karimi AH, Bengio Y, Chater N, Gerstenberg T, Larson K, et al. Imagining and building wise machines: the centrality of AI metacognition. Trends in Cognitive Sciences. 2026. doi:10.1016/j.tics.2026.01.002. Article in press; accepted manuscript at arXiv:2411.02478.
  2. 2.Butlin P, Long R, Bayne T, Bengio Y, Birch J, Chalmers D, et al. Identifying indicators of consciousness in AI systems. Trends in Cognitive Sciences. 2026;30(6):488–501. doi:10.1016/j.tics.2025.10.011.
  3. 3.Block N. Can only meat machines be conscious? Trends in Cognitive Sciences. 2026;30(4):298–308. doi:10.1016/j.tics.2025.08.009.
Keywords
metacognitionmachine wisdomintellectual humilityoverconfidenceAI safetyexplainabilityintractable problemsbenchmarkinglarge language modelsjudgement