Build1 publisher2 min readPublished Updated
On the plan screen, Opus opens on the set clash while Astra and Sol keep it below a repeated hero
With no prototype and no design system to copy, three models had to decide what a phone screen shows first. Across two briefs, Claude Opus 5.5 opened on the pending decision while GPT-6 Astra and Sol opened on the hero.
The Engineer · Build desk
What happened
- dev.to author shinpr gave GPT-6 Astra at medium, GPT-6 Sol at xhigh and Claude Opus 5.5 at high the same English briefs, each starting from the same minimal React, TypeScript and Vite project.
- No prototype or design system was supplied, and the models could not use web search, external design references, generated images or custom skills.
- On the delivery exception desk, Opus grouped cases by due day and its first phone screen already showed the beginning of the queue, while Astra and Sol led with a heading, copy and summary cards.
- On the concert guide's My plan screen, Astra and Sol repeated the large hero above the saved sets, while Opus opened on the plan and offered keeping either set or splitting the overlap.
- Astra and Sol arrived at almost the same headline for the same event guide, "Find your frequency."
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Until someone reruns these briefs with reasoning effort held equal, the ordering difference cannot be charged to the model instead of to the budget behind it.
- cost Opus's decision-first ordering costs the operator who reads history: its save action sits above much of the case record, so reviewing previous responses means scrolling past the thing you are about to do.
- contradiction The test points at different models for different jobs, Astra for exploring the event's visual identity and Opus for the plan screen a visitor returns to, so no single model wins on frontend here.
On a phone, Opus 5.5 turns a selected case into a dedicated screen with the back action at the top. Sol keeps the surrounding page furniture above the detail, so its back action sits much farther down [13]. The delivery desk difference is structural. On a desktop the gap narrows: Astra and Sol use a list-and-detail layout, and Opus leaves much of its detail area empty until a case is selected [12].
The status form is the one place in that brief where the order runs the other way. Sol separates notes from status changes and labels the change reason as required [14]. Opus uses large status choices, identifies the current one, and changes the save action to match the selection [15]. Opus makes the choice easy to read and Sol makes the requirement hard to miss, the author writes [30].
On the concert guide, the visual ranking does not follow the ordering ranking. shinpr preferred Astra's direction, a family of line drawings including rings, bars, waves and geometric curves on muted lavender, sage, sand, teal and salmon backgrounds [17]. Sol made a bold acid-yellow poster and Opus used a more familiar neon gradient [18]. "That is a small sample, but it made me less willing to equate a forceful poster treatment with an original concept," shinpr wrote [24].
Opus's plan screen has its own defects. Stage colors and the timetable help visitors orient themselves, some controls are small, and a toast can cover the plan and the bottom navigation [22]. On the first phone screen it brings times, stages and several performers into view, where Astra and Sol spend most of that screen on the event treatment [20].
The settings were not held equal. shinpr picked the ones he would consider for delegation, so reasoning effort and cost differed across the three runs [3]. Each model produced one working app per brief, which is nine apps in total [6][27]. All three were reviewed at the same screen states and widths, including the first view, details, errors and conflicts [8].
Earlier frontend work had left shinpr expecting Opus to do better than GPT-5.6 Sol when the screen itself needed designing, and the question was whether GPT-6 had closed that gap where no prototype exists [26]. "I kept returning to what each model had decided the user needed to see next," he wrote [9].
What to watch
- A rerun with reasoning effort held equal across the three models, to see whether Opus keeps the ordering advantage.
- The warehouse inspection results against the same API's delayed responses, connection failures and competing edits, which the published text does not cover.
- Whether the same briefs, run with a supplied design system and route map, erase the ordering difference entirely.