Build2 publishers2 min readPublished Updated
Cloudflare folds Iceberg ingest, catalog and SQL on R2 into one GA product called Basin
Cloudflare made its Data Platform generally available as Basin, three serverless services that ingest, catalog and query Apache Iceberg tables on R2. Outside Iceberg engines read the same tables, so Workers teams can price Basin SQL against a warehouse on identical data.
The Engineer · Build desk

What happened
- The three services were previously sold as Cloudflare Pipelines, R2 Data Catalog and R2 SQL, and now carry the Basin name.
- Basin Pipelines takes events from Workers, HTTP or Cloudflare Logpush, transforms them with SQL, and writes Iceberg tables or files into R2.
- During the open beta, Cloudflare's own billing and infrastructure teams used the services for billing-metric reporting and infrastructure telemetry.
- Cloudflare prices Basin on usage and attributes that billing model to the platform's serverless architecture.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A Workers team that already pays for a warehouse now has to choose whether Basin SQL replaces it, or Basin only stores and maintains tables the warehouse keeps querying.
- cost Zero-egress savings like Bobsled's reach teams whose data leaves storage across regions or clouds; a single-region, single-engine team will find most of any difference in compute.
- capability Basin Catalog runs compaction and planner statistics itself, so a team moving off self-managed Iceberg storage stops owning that maintenance job.
I think the query design is sound, because the component that rewrites the files is also the one that feeds the planner. Basin Catalog compacts metadata and data files to reduce I/O, and it generates statistics for query planning [6]. Basin SQL reads those statistics, splits a query into smaller tasks and spreads them across Workers [6]. Cloudflare says this keeps queries fast and consistent as datasets grow [6].
For new projects, Cloudflare says a Catalog, a Pipeline and a first Basin SQL query take seconds to set up. It aims that at coding agents that would otherwise wait and poll for resources or data [16].
Any Iceberg-compatible engine can read and write Basin data, according to Cloudflare, which names PyIceberg, DuckDB, Snowflake and Apache Spark [7]. The company says that portability is only possible with free egress [8]. Its stated reason for building Basin was that developers were already putting analytics data on R2, where the lack of egress charges made it practical to reach that data from different tools, teams, regions and cloud providers [15].
The customer quotes in the post describe other people's workloads. "Bobsled is a data product platform that the world's most advanced data teams use to build and distribute AI-ready data to partners, vendors and customers," said Julien Grobbelaar, head of platform at Bobsled [14]. "Basin allows us to build data products that can be made accessible in any region of every major data and AI platform, all at production-grade reliability and a fraction of the cost thanks to zero egress fees," he said [10]. Bobsled's product is data that leaves its storage. Grobbelaar credits the lower cost to zero egress fees [10].
Anomaly is the nearer comparison for a team building on Workers. "We moved our entire company's data pipeline to Basin Pipelines, Catalog, and SQL, replacing a complex AWS S3 and Athena setup with a cleaner, serverless architecture that reliably handles all of our event data," Dax Raad, Anomaly's co-founder, said in Cloudflare's post [9].
The available text of the announcement describes the billing model but stops before any per-unit rates for ingest, catalog maintenance or queries.
What to watch
- Whether Snowflake or Spark jobs writing to Basin tables alongside Basin Pipelines hold up under concurrent writes in production.
- Migration notes for R2 SQL and R2 Data Catalog beta users on whether bindings, endpoints or CLI commands change with the Basin names.
- Customer query timings at scale that test Cloudflare's claim that statistics-driven splitting keeps Basin SQL fast as tables grow.