Skip to content

Product1 publisher3 min readPublished

One food-delivery brief separated three flagship coding agents on design finish

XDA Developers handed Claude Code, Codex and Google Antigravity the same demanding site brief with no follow-ups. All three shipped every section that was asked for. The distance between them showed up in typography, spacing and mobile mockups.

The Product Desk · Product desk

Photograph accompanying One food-delivery brief separated three flagship coding agents on design finish
Photo: xda-developers.com

What happened

  • XDA Developers gave Claude Code, Codex and Google Antigravity the exact same prompt to rebuild a website, with no extra hints and no follow-up instructions during the first run.
  • The brief was a food-delivery site with a hero section, restaurant filters, responsive layouts, animations, testimonials and multiple interactive elements, with each agent making its own design decisions.
  • Codex, running GPT-5.6, produced what the reviewer called the best hero design of the three, layering ratings, delivery information and taglines around a large food image.
  • Gemini 3.8 in Antigravity impressed the reviewer on gradients and animation, while its FAQ section came out cramped and its mobile mockups looked artificial next to the rest of the page.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint An evaluation that asks whether the agent built the thing scores all three the same, because every defect in this run was finish rather than a missing feature.
  • cost The reviewer's standing habit of asking GPT-5.6 for bigger fonts is unpaid work that lands on whoever reviews the output, and it never shows up in the demo.
  • decision Teams shipping phone apps now have a reason to make mobile mockups part of the trial brief, since that is the one area where the reviewer's three candidates diverged most.
  • exposure Anyone quoting this comparison in a tooling decision is leaning on one person's taste on one brief, and owns that judgement internally.

The most useful line in XDA Developers' comparison is an aside about a habit. Working with GPT-5.6 on website design, the reviewer wrote, "I usually end up asking it to increase font sizes" [7]. That is a second prompt on work the tool had already handed back as done.

All three agents built the site. Codex kept every major section the brief asked for, with a sensible layout and food imagery that worked from the first run [5], and its hero section was the best of the three [6]. Claude Code was the one the reviewer described as complete from the start, with typography that felt right, consistent spacing, properly judged padding and intentional image placement [12].

The separation came from finish. Against Gemini 3.8, the reviewer named a cramped FAQ section and mobile mockups that "looked laughably artificial compared to the rest of the site" [10]. Against Codex, small fonts in secondary and supporting copy [7] and a newsletter section that "felt rushed, plain, and unfinished" [8]. Every defect reported is a finish problem; no section is missing and no build is broken [14]. XDA's reviewer wrote that the results were "nowhere near as close as I expected" [4], and what pulled them apart was the cleanup pass.

Mobile mockups are where the ranking actually turns. If your product is a phone app, one weak area in the generated preview is a rewrite you will do by hand. The reviewer wrote that both GPT-5.6 and Gemini fall apart on them, and that Claude Code's looked like believable app interfaces [13]. Antigravity's speed and its dark-gradient offer and newsletter sections earned it second place [11] [9].

This is three first-run outputs, one brief, one reviewer, and the article does not report scores [15]. Taste calls belong to the person making them. Nothing here tells you how the same agents behave on a six-year-old codebase, or inside your design system over a fortnight of tickets.

The version of this test you can run on Monday has two axes. First, completeness: did the agent produce every element of your real brief with no hand-holding. Second, rounds to ship: how many follow-up prompts, and whose time, it takes to get the output past your own review. Most credible candidates will tie on the first axis, which makes the second one the whole comparison. Log each fix as functional or taste while you do it, because the functional ones are the agent's problem and the taste ones are yours. The tradeoff in a one-brief run is that it rewards whichever agent's defaults happen to match your house style, so if your design system is fixed, run the brief twice with different components before you decide.

What to watch

  • A rerun that allows follow-up prompts would show how many cleanup rounds each tool needs to reach ship-ready.
  • Whether the next GPT-5.6 and Gemini 3.8 point releases fix the mobile mockup output both were marked down for.
  • Whether these tools change their default flagship agent, since the default is what most teams will actually run.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories