Build1 publisher3 min readPublished
Third-party trackers replace CMS's hospital price-file catalog after the data.cms.gov rebuild
Pennyforge found that CMS's Provider Data Catalog, the one federal list of hospital price-file URLs, is gone from data.cms.gov after the 2026-09-29 rebuild. Price pipelines now start from dated third-party trackers and depend on a few vendor hosts.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Pennyforge's checks found that CMS's new slug API returns 404 for the old catalog path and that the legacy SODA endpoint returns 410 Gone.
- Of the 6,583 distribution files in the rebuilt DCAT catalog, exactly one is a price file: the July-2026 enforcement CSV.
- A public GitHub tracker, cms-hpt-tracker, now lists all 5,419 CMS-listed hospitals and was checked row by row on 2026-09-29.
- The 3,977 distinct price-file URLs sit on 1,196 domains, and the top five of those domains serve 21.8% of them.
- In a 25-hospital spot check on 2026-10-02, five of the fetched price files were larger than 100 MB.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A crawl now starts from a third-party snapshot with its own check date, so the builder has to catch stale rows and coverage gaps between refreshes.
- exposure A single vendor's rate limit, auth change or outage can now cut off a large share of hospitals at once, so throttling and retries have to be set per host.
- cost Ingest code has to read each file's first bytes before choosing a parser, because both the extension and the Content-Type header misreport CSV files.
Pennyforge's case rests on several separate CMS surfaces showing the same result. The rebuilt site's 648-URL sitemap has no entry for the catalog. The Hospital Price Transparency program page now links only to the enforcement dataset and the GitHub docs repo [6]. The GitHub README has no catalog URL. The HPT FAQ, dated 2026-06-26, never mentions the catalog, so the documentation and the site do at least agree [15]. The post cites no CMS statement on whether the catalog was moved or retired [3].
The catalog gave a crawler every hospital's MRF URL in one machine-readable list [2]. Under the posting rule, each inpatient hospital must publish a machine-readable file of its charges, discounts and negotiated rates, plus a plain-text pointer file in its web root [1]. In Pennyforge's test, the pointer files that resolved returned three fields: location-name, source-page-url and mrf-url [16]. Without the catalog, a crawl runs in three steps. It takes a hospital list from a tracker, fetches each web root's pointer file, then follows the mrf-url. That is one more fetch per hospital, and one more thing that can fail. One sampled hospital had removed its root pointer file and returned 404 [14].
The files themselves vary more than their URLs suggest. The tracker's 2026-09-29 check covered 3,983 assessable hospitals with an MRF URL. About 68% of them serve a plain, statically named file, and nearly a quarter generate the file on demand [8]. File extensions undercount CSV. Four extensionless URLs in the sample were Azure Blob Storage endpoints serving plain CSV labelled application/octet-stream [9]. Counted by what is actually served, CSV is closer to two-thirds of the population [9]. Version strings need cleaning as well. Of the versioned listings, 3,334 (83.7%) are on template 3.0.0, about 36 are on 4.0.0 and about 200 are still on 2.x, with a long tail of malformed values [10].
A handful of vendors host much of the data. If the listed URLs were spread evenly, each domain would carry about 3.3 [1]. They are not spread evenly. The top five hosts are a Para Healthcare FS app, two ST Health Azure blob buckets, Hyve Healthcare's MRF host and Craneware's pricing API [12]. Azure Blob buckets alone carry 11.7% of listed files, and the top ten domains carry 29.7% [17]. The remaining 1,186 domains share the other 70.3% [2].
The spot-check figures hold only under the conditions Pennyforge tested in. The sample was 25 hospitals, seeded and random, fetched on 2026-10-02 from residential egress [13]. Seventeen returned 200 immediately, one redirected with a 302, one returned 403 because Cloudflare was blocking data-center IPs, and six timed out [14]. A production crawler running in a cloud region sends exactly the traffic that Cloudflare rule blocks, so I'd expect it to see more 403s. The size result is more useful for planning capacity than as a population rate. The large files were CSVs of 112 to 120 MB and a 106 MB JSON. The smallest was about 120 KB [13].
Pennyforge's census is careful work. Every figure comes with a dated probe, and the post is published under CC BY 4.0 [19]. Its own advice is to point pipelines at a dated, cited catalog and to expect the third-party rebuilds to drift [20]. Turquoise's vendor tracker, refreshed 2026-10-01, reports that 77.3% of hospitals meet both posting requirements [18]. For my own pipelines, I'd pin each crawl to one named tracker snapshot and its check date, and record which snapshot produced each row.
What to watch
- Whether CMS republishes hospital MRF URLs under a new data.cms.gov path or in the HPT GitHub docs repo.
- How often cms-hpt-tracker reruns its per-row check after 2026-09-29, since every crawl seeded from it inherits that date.
- How quickly hospitals move from template 3.0.0 to 4.0.0, which about 36 listings use now.