Science1 distinct publisher2 min readPublished
The winning condition both withheld solutions and made students redo the skill three times running, so what a district can copy from this experiment is a specification for how a tutor should behave rather than a shopping list.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Two design choices travelled together in the condition that worked. When a student got a problem wrong, NUMI walked them through the mistake instead of supplying the solution [3], and then required them to demonstrate the same skill correctly three times in a row [4]. That pair was measured against conventional computerized instruction [4], which means the three points cannot be apportioned between the withholding and the mastery gate [17]. Both are pedagogy, and only one of them is about answers.
The framing that reaches most readers is conditional: an AI tutor helps only if it holds back answers and coaches the student through the problem [16]. Oreopoulos grounds that in behaviour rather than technology, since students want to finish assignments quickly and, as he puts it, "easier equals less effort and less learning" [10]. But the reported experiment contains no condition in which an AI tutor does hand over the solution [18], so the conditional is a principle the results are consistent with rather than a contrast the study measured. Against the accumulating reports that AI harms learning by producing answers on demand [15], it is a reasonable principle to hold, and it remains a principle.
Three points is the number, and the account does not say three points out of how many, or how that compares with the spread of scores [5][6]. With more than 6,000 students in the sample [1], a difference that small can be statistically clean while sitting near the edge of what a teacher would notice, which is close to how Oreopoulos describes it when he refuses the game-changer language [8]. The write-up also does not say whether assignment happened at the student, classroom or school level [20], and in a school-delivered intervention that detail governs how much of the apparent precision survives clustering. Oreopoulos calls the work a proof of concept and possibly the first objective evidence that a well-designed AI tutor beats instruction without AI [7]; the paper sits alongside two other recent working papers from the same line of research at the National Bureau of Economic Research [11].
His choice of comparison is the part worth carrying forward. He does not claim NUMI beats a good human tutor. He says a good human tutor is usually better, and that the question worth asking is what a well-designed AI tutor does for a student who is stuck with nobody available to help [9]. That is a defensible baseline, and it is the one in which a modest gain counts for something. It is also the baseline that makes the mastery loop, not the model, the object anyone should be trying to reproduce.
Ranked by verification strength, evidence, and original report placement.
A team led by University of Toronto Mississauga economics professor Philip Oreopoulos conducted an experiment involving more than 6,000 middle school math students in Tennessee to determine under what conditions AI tutors could be beneficial.
Oreopoulos collaborated with Michael Liut, an associate professor of mathematical and computational sciences at UTM, to build NUMI, an AI tutor designed to avoid giving direct math solutions and instead provide structured support to students.
In the study, NUMI taught students how to solve math problems by reviewing past mistakes while withholding direct solutions, with support after a mistake designed to promote step-by-step mathematical reasoning rather than answer-getting.
Students scored higher on a test covering practiced and unpracticed math problems when an AI tutor walked them through their mistakes and then made them demonstrate the same skill correctly three times in a row, compared with students receiving conventional computerized instruction.
Students assigned to NUMI gained what the report calls a modest learning advantage, scoring three points higher than those without an AI tutor.
The phys.org report gives the size of the gain as three points but does not state the test's total scale or the standard deviation of scores.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
SNAP soda bans cut purchases 12%. The exclusion list decides what that is worth1 distinct publisher
product
Eight suppliers, 53 bids, £300,000 each: Britain turned AI tutoring into a procurement line1 distinct publisher
science
An AI-drafted pre-bunk still blunted rumor damage to voter confidence a week later1 distinct publisher
build
Students using GPT-4o scored nearly a full point higher on Bocconi's five-point grading scale1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Large trial, unreadable number
The trial is the strong part: over 6,000 students, a baseline of the computerized instruction schools actually use, a test spanning practiced and unpracticed problems, and a working paper with a DOI anyone can pull. The reporting is the weak part. Three points out of how many, with what spread, and assigned at what level — student, classroom, school — are all absent, and a reader cannot tell whether this is a rounding error or a meaningful shift. One account, sourced to the team that built the tutor, and no one has yet read the paper back at them.
Researcher-arranged use
Six thousand Tennessee middle schoolers is real classroom usage, but it is usage the researchers arranged and controlled. Nothing in this reporting shows a district choosing NUMI, a licence, a price, a second state, or any life for the tutor outside the experiment that produced it. What exists beyond the trial is a working paper.
Overshoots at the top, sober in the middle
Unusually, the researcher is the brake and the framing is the accelerator. Oreopoulos rules out 'game changer', calls the advantage modest, and concedes a good human tutor usually wins. The overreach is structural: the 'only if it holds back answers' conditional needs an answer-giving tutor to have been tested, and none was, while an unnamed body of evidence that AI harms learning is invoked to make three uninterpretable points feel consequential.
Builders reporting on their build
The tutor's authors are also the study's authors and the story's only quoted voices, and their stated next move is to build a more engaging version — so a positive first result is what unlocks the follow-on work. It reaches readers through a university communications channel picked up by phys.org, with no funder disclosed, no vendor relationship described, and no independent evaluator anywhere in the piece. That said, the same interested party is the one saying human tutors are usually better, which is not how a promotional account behaves.
Single-channel, unresolved units
We can be fairly sure what was claimed and by whom; the paper is identified down to its DOI. We cannot yet say how large the effect is, because the units are missing, nor whether the design supports the conditional the story leads with. One publisher, one interested set of sources — enough to report, not enough to bank.