Science1 distinct publisher3 min readUpdated
On Quanta's podcast, the Santa Fe Institute's Melanie Mitchell says we lack adequate methods for measuring machine cognition, and points to how psychologists test babies and animals.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
Melanie Mitchell, a cognitive scientist and computer scientist at the Santa Fe Institute, told Quanta Magazine's podcast The Joy of Why that the field lacks adequate methods for measuring machine cognition, and that AI is better described as a form of "alien intelligence" that operates through non-human cognitive mechanisms [2][3][4]. Her proposed remedy is not another leaderboard: adapt the methods psychologists already use on the two alien intelligences they cannot interview, babies and animals [5].
The stake is operational, not philosophical. Quanta frames the question as whether a language model answering a question is reasoning like a human or producing text that looks like reasoning [1], and notes that the answer determines what we can trust these systems to do, how closely they need supervising, and what their real-world impact turns out to be [7].
Follow that framing to its consequence, and a lot of current benchmark reporting stops meaning what it says. If two different mechanisms can produce the same score on a test, the test does not tell you which mechanism produced it [14]. That is the whole discipline of infant and animal cognition work: you cannot ask a crow to explain itself, so you design a control condition where the ability you are claiming is absent but every other route to the right answer is still open. Benchmarks built as human exams do not carry that control. They were written for subjects who share our failure modes, and they are scored as if a pass licenses a claim about internal process. Mitchell does not use the word unfalsifiable in the material we have; that inference is ours.
The episode's cautionary tale is a horse from the early 1900s that appeared to do arithmetic, offered as a warning about how we assess intelligence [8]. Quanta's introduction does not name the animal [15]. The point survives the omission: the performance was real and the attributed mechanism was wrong, which is exactly the error a score-only evaluation cannot catch.
Co-host Janna Levin put the epistemic problem plainly, saying we are excited about the artificial mind while having very little comprehension of the human mind, and that we are trying to skip a step [11]. Steven Strogatz pointed to comparative psychology, the study of intelligence in birds, dogs and dolphins, as the neglected reference class [12]. Mitchell had appeared on the show roughly five years earlier, before ChatGPT arrived around November 2022 [13]; the hosts date this recording to July 23, 2026 [10], which puts it about 44 months into the era of widely deployed language models [16].
Mitchell lays out six principles for better assessing machine cognition [6]. The Quanta introduction announces them without enumerating them, and the transcript portion available to us covers only the opening setup [15], so the substance of the proposal is in the audio rather than in anything quotable here.
What to watch: whether those six principles include control conditions of the kind animal work requires, and whether any lab reports them alongside a benchmark number. Watch also how interpretability is positioned, since the episode treats the difficulty of reading what happens inside these systems as an open problem rather than a solved check [8], and whether the AI-assisted mathematics results discussed in the same conversation [8] get evaluated as capability or as process.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Mitchell tells Steven Strogatz that methods psychologists use to study cognition in other kinds of "alien intelligence" - babies and animals - can be adapted to probe AI.
Mitchell lays out six principles for better assessing machine cognition.
Strogatz says in the recording: "As we speak, it's July 23rd, 2026."
Strogatz points to comparative psychology, where researchers look at intelligence in birds, dogs or dolphins, and says we have a lot to learn about thinking about intelligences other than our own adult human intelligence.
Mitchell appeared on the show previously, about five years earlier and before ChatGPT; Strogatz says the earlier conversation was maybe in 2021 and that the ChatGPT wave hit around November 2022, which Mitchell confirms.
The Quanta introduction states that Mitchell lays out six principles and refers to a math-performing horse from the early 1900s without naming the animal or enumerating the principles; the transcript text supplied covers only the hosts' opening setup and the first exchange with Mitchell.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source attributed argument, primary transcript, no data
Provenance is good for what was said — a verbatim, dated Quanta transcript with named speakers — but the substance is an expert's argument rather than measurement. No benchmark result, study, dataset or third-party corroboration is supplied, the six announced principles are not enumerated in the available text, and only one publisher covers the story.
No adoption evidence in supplied sources
The source is a podcast introduction and transcript excerpt about how to assess machine cognition. It reports no releases, deployments, benchmark runs, pricing or license changes, usage disclosures or incidents, and no third party is shown adopting Mitchell's proposed evaluation approach, so adoption cannot be scored.
Mildly overstated packaging around modest claims
The argument itself is stated cautiously and hedged as one researcher's view, which is well matched to the evidence. The small positive gap comes from packaging: the introduction advertises six principles, AI-assisted mathematics breakthroughs and a cautionary historical anecdote that the supplied text never delivers or substantiates, so the promised payload exceeds what is shown.
Publisher covering its own podcast; no vendor stake visible
The clearest incentive is self-distribution: Quanta Magazine is publishing an article about The Joy of Why, its own podcast, and directs readers to Apple Podcasts, Spotify, TuneIn or Quanta's stream. No commercial AI vendor, funder relationship or product being sold is disclosed in the supplied text, and the guest's stake is limited to her own research program, so the incentive load is real but modest.
Moderate: reliable quotes, thin substantiation, no corroboration
Confidence is high that the quoted statements were made — the transcript is verbatim, dated and attributed — but low that the underlying thesis has been tested here. One publisher, one item, a truncated excerpt, unenumerated principles and no adoption evidence hold the overall figure below the midpoint.
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
invest
A Connecticut judge just priced prompt injection: no fine, no e-filing2 distinct publishers
science
A vaccine communicator's testable claim: hesitancy answers to burden, not to better facts1 distinct publisher
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 20, 2026