Science1 distinct publisher3 min readPublished
The panel's argument is a measurement one: clicks and streaks track attention rather than learning, and generative AI widens the gap by improving a student's output while leaving the underlying skill where it was.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Engagement wins as a metric because it is the metric the software produces for nothing. A learning platform logs session length and reward streaks as a byproduct of running; establishing that a 12-year-old can use a concept later, on paper, in a setting the vendor did not build, requires someone to design an assessment, administer it away from the product, and find students to compare against. That asymmetry is why dashboards became the evidence base for purchasing. The APA report declines to accept the cheap measure, arguing that applying knowledge outside the app is the much stronger indication that learning happened, and that frequent clicking or responding to animations does not necessarily mean it did [3][5].
Generative AI makes the gap wider by acting on the observable side of it. "AI can help a student produce a better essay, but the improved final product does not necessarily mean the student has become a better writer," said Nicole Barnes, the APA's executive lead psychologist for education, who added that schools will increasingly need to separate performance from learning as well as engagement from learning [9]. The report frames this as a construct validity problem, not a cheating one. The artifact a teacher grades and the capability a school is trying to build have come apart, and the report's stated risk is precisely that immediate performance rises without underlying knowledge or skill following [2].
Rather than call for exclusion, the panel recommends teaching students to interrogate what a chatbot returns: question its assumptions, look for gaps in its reasoning, compare it against other sources, and ask how the answer would change under different circumstances [10].
The procurement half is where this bites. The report places primary responsibility for evidence on the companies that build these products and on the institutions that adopt them, tells developers to be ready to demonstrate that their products improve learning, and tells schools and districts to expect credible, independent evidence before they spend [11]. APA chief executive Arthur C. Evans Jr. framed the recommendations as a way to determine whether tools actually produce the outcomes they promise [13].
What the summary leaves undefined is where that bar sits. The published summary sets out 10 recommendations without effect sizes, and without defining what makes evidence credible or independent [7][15]. Independent of whom, funded by whom, with transfer measured how long after use, against which comparison group: a district writing that phrase into a contract has to supply those numbers itself, and the guidance covers roughly 13 school years, ages 5 to 18, in which the right assessment looks different every couple of grades [4][14].
For parents the version is cruder and probably more usable. Spend a minute watching the screen; if the experience is largely dressing a character, collecting rewards or responding to flashing images, question the educational value, then ask a teacher or school board whether independent evidence exists that the app leads to learning [12]. Note the report's own hedge, which is doing real work: engagement does not *necessarily* mean learning [5]. The panel treats attention as a precondition it never dismisses, not as the outcome it scores.
Ranked by verification strength, evidence, and original report placement.
The American Psychological Association published the APA Expert Report on Children's and Adolescents' Learning with Educational Technology, warning that student engagement with educational technology is not the same as learning.
The report warns that generative AI tools pose particular risks by improving students' immediate performance without building their underlying knowledge or skills.
The report argues that students' ability to apply new knowledge outside an app or screen is a much stronger indication of learning than how much time a child spends on a screen.
The report addresses parents and educators working with children and young people ages 5 to 18, and says they should be wary of marketing claims that run ahead of scientific evidence.
The report states that even if a student is paying close attention to an app, by clicking frequently, responding to animations or interacting with rewards, that engagement does not necessarily mean the student is learning.
The report was compiled by a multidisciplinary panel of cognitive and educational psychologists, learning scientists and EdTech experts, and synthesizes evidence across educational apps, games, intelligent tutoring systems, generative AI tools and digital learning platforms.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
A school agenda shipped with "Vitoiis" and a planet named Marc, and no one read it first1 distinct publisher
invest
California's SB 903 would put a clinician in front of every mental health chatbot1 distinct publisher
product
Schools swapped AI bans for supervised failure, and the cost moved to teacher training1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One faithful relay of the association's own announcement
Every load-carrying line — the ten recommendations, the panel's makeup, the sentence putting the burden on developers and districts — comes from a single Phys.org write-up that follows the APA's release closely. The quotes are on the record and the positions are unambiguous, which is worth something. But nobody outside the association has read the underlying synthesis for us, and the summary names no study, no effect size, and no test for when knowledge counts as having transferred.
Nothing bought, blocked or rewritten yet
Recommendations are not uptake. Our reporting shows no district that has changed a purchasing rule, no vendor publishing the learning evidence the report asks for, and no policymaker citing it. Until one of those appears there is nothing here to measure, and inventing a number would be exactly the error the report warns about.
Sound diagnosis, unenforced remedy
The panel's own claims are unusually restrained — Evans opens by calling edtech genuinely promising, and Barnes's point about essays versus writers is a careful distinction, not a scare. The overstatement sits in the word 'shifts'. 'Developers should be prepared to demonstrate' is a norm with no metric attached and no one assigned to check it, and nothing in this reporting shows a buyer acting on it. As an argument about measurement it holds; as a transfer of burden it is so far a wish.
The yardstick belongs to the panel that proposed it
The APA is arguing that psychological science should be the measure of whether edtech works, and one of its ten recommendations asks for more investment in that research — a predictable ask from the body that convened the panel, and a legitimate one worth naming out loud. On the other side, the companies told to prove their products work are absent from the piece entirely, and Phys.org adds no scrutiny of either interest.
Positions clear, consequences unknown
We can be fairly sure what the APA said and who said it; the quotes and the recommendation count are concrete. What we cannot judge from a single announcement-based story is whether the report has any force — whether the evidence it synthesises is strong, whether districts will act, whether developers will comply. High confidence in the words, low confidence in the effect.