Product1 distinct publisher3 min readPublished
Four researchers quoted by The Verge trace the noodly, hole-riddled AI menu shot back to how diffusion builds an image, which makes the ugliness a reproducible failure mode rather than an unlucky prompt.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The person who has to approve a menu photo can use Giovanbattista Califano's sentence as a sorting rule. Add to his strand problem the repeating textures he also names, bubbles and seeds, which he says diffusion models struggle to keep inside sensible boundaries so they spill into places they should not be [6]. That gives five named shapes [10], short enough to check a menu against before anyone opens a tool: a bowl of pho and a plate of shredded chicken are mostly strands, a sesame bun and a poppy seed loaf are mostly repeating dots, a whole roast or a slab of cake is one continuous mass.
Teams often respond to a bad AI food image by rewriting the prompt with more specific language. The pipeline Chris Russell of Oxford describes puts the error before the adjectives arrive. Diffusion begins with pure noise and removes it gradually, so coarse structure is recovered first and fine texture detail comes at the end [3]. More descriptive wording buys better texture, but the coarse structure is already locked in before that stage, and that earlier stage is where the failure actually happens. This is not one vendor's quirk either; many of the leading image generators work this way [11].
Michael Cook of King's College London supplies the other half of the problem: the models have no understanding of why food should not look like other, non-food images, which is why a dessert can come out reading as masonry [8].
The gap in the reporting is the price tag. None of the quoted researchers puts a number on what an unappetizing AI image costs a cafe in covers or a brand in sales, and Califano's own field is human response to AI-generated imagery [5]. So the internal case for a camera rests on avoidable defect rates, not on a measured revenue hit. That is a weaker argument than a conversion chart, and it is the one the evidence actually supports.
The grid worth drawing has two axes: whether the dish's silhouette depends on thin strands or repeated small features, and whether the food itself is within reach of a camera. Strand-heavy and in reach, which describes most restaurant work, is a shoot, because the failure lives in the subject and no amount of review will argue it out. Blocky and in reach is also a shoot, since the marginal cost of a phone and a window is lower than a third round of sign-off on an image that still looks wrong. Blocky and out of reach is where generation is defensible, with the edges and any seeded surfaces checked before publication. Strand-heavy and out of reach is where the honest answer is illustration, or no image at all.
The useful shift is from taste to geometry. A reviewer who says "this looks off" loses the argument to whoever made it, but naming the dish as tendrils and seeds, the documented failure case, describes a defect class, and defect classes can be tested for on a Monday rather than defended on a Friday.
Ranked by verification strength, evidence, and original report placement.
The Verge reports that restaurants, cafes and brands are increasingly turning to AI to generate images promoting their food, producing a torrent of unappetizing output.
Examples catalogued by The Verge include donut shrimp, wormlike noodles, stringy chicken, ice cream resembling cracked concrete, burgers seemingly fashioned from rocks, and clustered holes of the kind that can trigger trypophobia.
Chris Russell, a professor of AI, government and policy at the University of Oxford and a computer vision expert, said diffusion models start from an image of pure noise and gradually remove noise, so initially coarse structures are recovered first with fine texture details coming at the end.
Russell said that in many unsettling food images the model gets the basic structure of an object wrong at an earlier stage and then lumps vivid texture details on top of it, the same kind of failure as a person generated with six fingers instead of five.
Giovanbattista Califano, a behavioral scientist at the University of Naples Federico II who studies responses to AI-generated imagery, said: "Diffusion models are notoriously weak at generating thin, continuous, terminating structures. Noodles, strands, and tendrils are exactly the kind of geometry that trips this up, so you get spaghetti-like artifacts bleeding into places with no anatomical or culinary logic."
Califano added that other repeating textures such as bubbles and seeds are similarly hard for diffusion models to contain within sensible boundaries, so they often spill into areas they should not be in.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Oxford's Zostavax-to-Shingrix switch closes the healthy-vaccinee loophole, and opens another2 distinct publishers
product
Hochul defends a Meta teen settlement whose age checks land on every adult account1 distinct publisher
build
The FCC read HOVERAir's Versa by what it was built to do, not how it was boxed2 distinct publishers
science
Northern Ireland's phone-pouch pilot changed classroom behaviour faster than pupil well-being1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Four named academics, one newsroom
Oxford, Naples Federico II, Zurich and King's College London are all on the record with direct quotes and stated specialties — more rigour than this genre usually bothers with, and two of the four describe the same coarse-to-fine pipeline from different starting points. What is missing is anything testable: no generator is named or versioned, no offending image is traced back to the model that made it, and nobody reproduced the spaghetti artifact on demand. This is expert reasoning, not measurement.
Screenshots, no denominator
Not one restaurant, chain, agency or tool is named, and no count, spend figure or platform datapoint appears. The images are real, but how many businesses have swapped a photographer for a prompt is simply not on the record in this reporting, so there is nothing here to score.
Slight reach beyond the interviews
The core move — treating the ugliness as a predictable property of how diffusion builds an image rather than an unlucky prompt — is a genuine explanatory gain, and The Verge states it without inflation. It does run a little ahead of its own evidence: four conversations establish why thin, terminating shapes are hard, but nothing shown here demonstrates the failure recurring across the generators actually in use. Mild overreach, not hype.
Nobody in the room is selling a model
The four voices are academics, and no lab, vendor or brand appears to promote or defend a generator, so commercial pressure on the substance is low. The pull that does exist is editorial: a piece assembled from the most grotesque examples rewards extremes, and The Verge happily leans into the bit about the burrito from hell. Light pressure on the claims, heavier pressure on the picture selection.
Coherent diagnosis, unmeasured scale
One newsroom, four sources of its own choosing, no corroborating outlet, nothing quantified — the ceiling here comes from structure rather than sloppiness. The mechanism holds together well enough that I would bet on the diagnosis of why noodles and bubble textures go wrong; the size of the phenomenon and its persistence in current models stay open.