Skip to content

Build1 publisher3 min readPublished

Third-party trackers replace CMS's hospital price-file catalog after the data.cms.gov rebuild

Pennyforge found that CMS's Provider Data Catalog, the one federal list of hospital price-file URLs, is gone from data.cms.gov after the 2026-09-29 rebuild. Price pipelines now start from dated third-party trackers and depend on a few vendor hosts.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Third-party trackers replace CMS's hospital price-file catalog after the data.cms.gov rebuild
Generated illustration

What happened

  • Pennyforge's checks found that CMS's new slug API returns 404 for the old catalog path and that the legacy SODA endpoint returns 410 Gone.
  • Of the 6,583 distribution files in the rebuilt DCAT catalog, exactly one is a price file: the July-2026 enforcement CSV.
  • A public GitHub tracker, cms-hpt-tracker, now lists all 5,419 CMS-listed hospitals and was checked row by row on 2026-09-29.
  • The 3,977 distinct price-file URLs sit on 1,196 domains, and the top five of those domains serve 21.8% of them.
  • In a 25-hospital spot check on 2026-10-02, five of the fetched price files were larger than 100 MB.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A crawl now starts from a third-party snapshot with its own check date, so the builder has to catch stale rows and coverage gaps between refreshes.
  • exposure A single vendor's rate limit, auth change or outage can now cut off a large share of hospitals at once, so throttling and retries have to be set per host.
  • cost Ingest code has to read each file's first bytes before choosing a parser, because both the extension and the Content-Type header misreport CSV files.

Pennyforge's case rests on several separate CMS surfaces showing the same result. The rebuilt site's 648-URL sitemap has no entry for the catalog. The Hospital Price Transparency program page now links only to the enforcement dataset and the GitHub docs repo [6]. The GitHub README has no catalog URL. The HPT FAQ, dated 2026-06-26, never mentions the catalog, so the documentation and the site do at least agree [15]. The post cites no CMS statement on whether the catalog was moved or retired [3].

The catalog gave a crawler every hospital's MRF URL in one machine-readable list [2]. Under the posting rule, each inpatient hospital must publish a machine-readable file of its charges, discounts and negotiated rates, plus a plain-text pointer file in its web root [1]. In Pennyforge's test, the pointer files that resolved returned three fields: location-name, source-page-url and mrf-url [16]. Without the catalog, a crawl runs in three steps. It takes a hospital list from a tracker, fetches each web root's pointer file, then follows the mrf-url. That is one more fetch per hospital, and one more thing that can fail. One sampled hospital had removed its root pointer file and returned 404 [14].

The files themselves vary more than their URLs suggest. The tracker's 2026-09-29 check covered 3,983 assessable hospitals with an MRF URL. About 68% of them serve a plain, statically named file, and nearly a quarter generate the file on demand [8]. File extensions undercount CSV. Four extensionless URLs in the sample were Azure Blob Storage endpoints serving plain CSV labelled application/octet-stream [9]. Counted by what is actually served, CSV is closer to two-thirds of the population [9]. Version strings need cleaning as well. Of the versioned listings, 3,334 (83.7%) are on template 3.0.0, about 36 are on 4.0.0 and about 200 are still on 2.x, with a long tail of malformed values [10].

A handful of vendors host much of the data. If the listed URLs were spread evenly, each domain would carry about 3.3 [1]. They are not spread evenly. The top five hosts are a Para Healthcare FS app, two ST Health Azure blob buckets, Hyve Healthcare's MRF host and Craneware's pricing API [12]. Azure Blob buckets alone carry 11.7% of listed files, and the top ten domains carry 29.7% [17]. The remaining 1,186 domains share the other 70.3% [2].

The spot-check figures hold only under the conditions Pennyforge tested in. The sample was 25 hospitals, seeded and random, fetched on 2026-10-02 from residential egress [13]. Seventeen returned 200 immediately, one redirected with a 302, one returned 403 because Cloudflare was blocking data-center IPs, and six timed out [14]. A production crawler running in a cloud region sends exactly the traffic that Cloudflare rule blocks, so I'd expect it to see more 403s. The size result is more useful for planning capacity than as a population rate. The large files were CSVs of 112 to 120 MB and a 106 MB JSON. The smallest was about 120 KB [13].

Pennyforge's census is careful work. Every figure comes with a dated probe, and the post is published under CC BY 4.0 [19]. Its own advice is to point pipelines at a dated, cited catalog and to expect the third-party rebuilds to drift [20]. Turquoise's vendor tracker, refreshed 2026-10-01, reports that 77.3% of hospitals meet both posting requirements [18]. For my own pipelines, I'd pin each crawl to one named tracker snapshot and its check date, and record which snapshot produced each row.

What to watch

  • Whether CMS republishes hospital MRF URLs under a new data.cms.gov path or in the HPT GitHub docs repo.
  • How often cms-hpt-tracker reruns its per-row check after 2026-09-29, since every crawl seeded from it inherits that date.
  • How quickly hospitals move from template 3.0.0 to 4.0.0, which about 36 listings use now.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories