Build1 distinct publisher3 min readPublished
A randomized trial with 1,053 freshmen graded 180-word marketing memos, and the traits that lifted scores were the ones a model supplies cheaply, while falsifiability and divergence from classmates cost points.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
The Tokenizer Is Your Real Price List, Not the Per-Million Rate Card1 distinct publisher
product
Rillet's $100M reads as proof mid-market ERP is rip-and-replace, mostly at the cheap end1 distinct publisher
leadership
Disney swaps raises for discounted stock and a full health-plan re-enrollment1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
Which rubric terms carried a negative sign is the part worth reading twice. Within a single answer, idea diversity correlated with higher scores [12]. Divergence from what other students wrote correlated with lower ones [13]. Those two measure spread against different baselines: one inside the text, one against the class. A scheme that pays for the first and charges for the second is scoring proximity to an expected solution space, which is what the authors say this assignment's score mainly rewarded [14].
That is the shape of output a general-purpose model produces without being asked. The GPT-4o arm carried about two more ideas per answer, read as more logically coherent, and landed closer to what three subject-matter experts recommended [6]. Nearly a full point on a 1-to-5 scale is roughly a quarter of the usable range [5][1]. The gap survived controls for argumentation quality, idea count, idea diversity and text properties [7], and the authors read the remainder as content quality rather than student knowledge [8]. Mechanically, that residual means the named features did not account for the whole effect, so graders were responding to something the measured feature set never named.
For the number to move to your hiring exercise, the exercise has to reward the same things. The task here was up to 180 words of marketing recommendations for the university's own merchandise shop [3]: short, structured, aimed at an answer someone already knows. If your take-home instead asks how a proposal fails, note that stronger falsifiability and more detailed explanation of how an action would work both correlated with lower scores here [13], and the causal-reasoning arm's slightly lower average is what that tax looked like in practice [9]. Stacking GPT-4o on top of that lesson bought no further gain on traditional scores [11].
The design limits deserve plain statement. Randomization ran across 13 sections rather than individually among all 1,053 students [19], which works out to about 81 students per section [2] and leaves a lot fewer independent coin flips than the headline count implies. The primary score came from human graders blind to condition [20], while the secondary constructs, including causal reasoning and idea diversity, were scored by models from OpenAI and Anthropic [21]. OpenAI supplies the technology under study and was involved in the research, with several authors employed there during it [22]. There was no later session without ChatGPT, and the authors acknowledge they cannot say whether the gain came from knowledge students gained [17], so what is demonstrated is graded performance on one assignment [18].
The cheapest change here is a rubric line rather than a prohibition. The authors' own conclusion is that diversity and originality have to be written into the criteria explicitly if they are meant to count [15]. Scoring "what would have to be true for this to work, and how would we notice it failing" costs a grader one extra pass, and it prices the part of the work a 180-word polish engine does not supply. The-decoder's account of the study puts the underlying problem more bluntly: current grading measures polish, structure and completeness rather than learning and understanding, which is what makes AI easy to use as a cheating tool [16].
Ranked by verification strength, evidence, and original report placement.
A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment.
In November 2025, 13 sections of an introductory management course were randomly split into four groups: control, causal reasoning lesson, GPT-4o access, or both.
Students had to write marketing recommendations for the university's merchandise shop in up to 180 words.
The lesson covered coherent causal logic, falsifiability, and how a proposed action might lead to a desired outcome.
Students with GPT-4o scored nearly a full point higher on a 1-to-5 scale.
Answers from the GPT-4o group contained about two more ideas on average, were more logically coherent, and more closely matched the recommendations of three subject-matter experts.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 30, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Clean design, single unverified account
The design is unusually tidy for education research — random assignment across 13 sections, graders blind to condition, controls for idea count, argument quality and text properties, and a residual effect that survives them. What our coverage cannot do is check the paper: no citation, no venue, no peer-review status, no replication, and the softer outcomes were scored by models from the two labs with the most riding on the answer.
Trial-scale use, no rubric has moved
One management course, one Milan business school, one month — that is the whole of the direct usage. The prior work The Decoder cites shows students already reaching for these tools at scale, half a million US grades and 26,000 students in China, but nobody in this reporting has changed a rubric, a syllabus or a procurement decision because of the Bocconi result.
The number outruns its footnotes
Nearly a full point is a real effect, but it was earned on 180-word memos scored by a rubric the same study shows rewards convention, in an experiment that randomized sections rather than students, by authors who admit they never checked what students could do once the model was taken away. The Decoder prints all of that; the gap is what happens to the figure when it travels without those lines attached.
Vendor, co-author and scorer in one chain
OpenAI supplies the model, took part in the research, and — as The Decoder notes — employed several of the authors during it, then framed the finding as a challenge for grading rather than a question about learning. The measurement chain compounds the exposure: causal reasoning and idea diversity, the outcomes that need judgement, were scored by OpenAI's and Anthropic's own models. None of this makes the effect fake; it does mean nearly every step from intervention to outcome passed through an interested party.
Half-confident, and for stated reasons
All of the above rests on a single account of a study we cannot inspect, so our confidence sits mid-scale. The blinded primary outcome and the frankly stated limitations pull it up; the absent paper citation, the absent second outlet and the absent post-treatment test pull it back.