Build1 publisher3 min readPublished
Shopify's products.json endpoint fed three quarters of a shopping agent's finds across 85 fashion sources
Twenty-one Shopify stores supplied 3,173 of the 4,112 sale items a homelab shopping agent pulled from 85 fashion sources in one run. What each store returned depended on whether its own front end already served structured prices and per-size stock.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Only 45 of the 85 sources returned even one item on the measured run.
- After filtering, 941 items were confirmed in the author's size and 452 could not be size-checked, with wrong size the largest single cut at 1,411.
- Eight sources, among them Nike, Puma, Selfridges and Clarks, list prices without per-size stock the agent can read, so their items are tagged size_unknown and kept out of alerts.
- A large retailer's public, search-only Algolia key, taken from its own sale page, returned 400 items already filtered server-side on the same run.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Building an agent that reads retail sites now starts with inspecting the requests a store's own front end makes; an HTML parser is the fallback for stores that expose nothing structured.
- constraint A retailer that loads inventory in a later call, or shows one size selector for every stock state, limits what a notifier can promise, and wider coverage then costs alert precision.
- exposure A harness that turns an unreadable robots.txt into 'disallowed' misstates which retailers refuse automated access, and the error persists because the outcome often looks correct.
Any Shopify storefront answers at /products.json with paginated JSON, no key and no rendering [8]. The response lists every product and every variant, each with `price` and `compare_at_price` [8]. To check whether an item is really reduced, you subtract one field from the other. Nobody has to guess at strikethrough markup [8]. The variants also carry per-size availability. "That one field is the whole reason this project works," the author wrote in a dev.to post [9].
Per source, the gap is wide. The 21 Shopify stores averaged about 151 items each [3]. The other 64 sources produced 939 items between them, under 15 apiece [2][3]. On this run, a Shopify store yielded roughly ten times what an average source outside that group did [3].
I think the platform is a proxy. The real variable is whether the store's own front end already requests something structured. The Algolia retailer gets the same result through a different interface. Its server-filtered result was more than two and a half times the Shopify per-store average, and it arrived already cut down to sale menswear in the author's sizes [10][6]. The failures split along the same line. Puma loads its size grid and inventory from a later API call, so none of it is in the HTML [12]. On other sites the size selector looks the same whether a size is in stock or gone [12]. The author's rule is to "check whether the shop's own front end is already calling something cleaner than the HTML" before writing a scraper [20].
The harness logic still mattered, mostly in how it labelled failure. Five well-known retailers were logged as robots.txt refusals. On those sites the file sat behind the same edge protection as everything else and returned 403 to a plain HTTP request. The code could not parse any rules and fell through to "assume disallowed" [15]. Fetched the way a browser fetches it, every one of the five files allows `*` [15]. "Failing to read a policy is not the same as the policy saying no," the author wrote [16]. Re-tested properly, all five still returned nothing: missing men's prices on one, redirects and a region picker on two, empty pages on two [17]. The author summed it up as "Right answer, wrong reason, for about a day" [21].
Rendering changed the result in two cases. New sources were probed twice, once with a plain HTTP client and once in a real browser, with the same User-Agent, the same robots check and no evasion [18]. A plain client alone would have filed two of them as blocked. One rendered its product grid in JavaScript, and the other would not answer a bare client but served a browser [18]. Sports Direct and a high-street chain returned 404 to every URL tried, because the author was guessing their category paths wrong [19].
The three-quarter share describes one list on one day. It comes from the 19 September gather, filtered to one person's menswear sizes [2][6][1]. The post does not say how the 85 sources were chosen. For the share to hold elsewhere, another agent's source list would need a similar weight of Shopify stores. The data that agent needs would also have to sit in structured fields like per-variant stock. Eight sources in this run did not expose that field in a form the agent could read [11].
What to watch
- Any change by Shopify to the default public /products.json response, such as a key requirement, would remove the source of most of this agent's items.
- A repeat run on another date, or against a different source list, would show whether Shopify's share stays above three quarters.
- Whether the retailer that exposes a search-only Algolia key on its sale page keeps that key public.