Build1 publisher2 min readPublished
Aurora's aurora_analytics extension prints DuckDB operators in its EXPLAIN output
A dev.to walkthrough found DuckDB's operator names in the EXPLAIN output of AWS's aurora_analytics extension, on a Parquet query that executed in 81 ms. That run hit a warm cache on one 500,000-row file, so it tells you which engine ran far more than how fast it will be on lake-scale data.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- AWS announced on September 30 that Aurora PostgreSQL can run queries over Apache Iceberg and Parquet files in data lakes through an embedded DuckDB engine, next to its operational data.
- The author tested the claim by pulling an execution plan from Aurora and one from open-source pg_duckdb for the same query over the same Parquet file.
- A foreign table created with an empty column list makes Aurora infer the schema from the Parquet file's footer.
- Amazon agreed in August 2026 to acquire DuckLabs, the company behind DuckDB, while the DuckDB open-source project remains independent.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Before a team retires a separate lake engine or ETL step, it can diff Aurora's operator tree against pg_duckdb's on its own queries and confirm which engine's behavior it is taking on.
- cost The database team pays for setup up front, with a version floor, a parameter group flag, IAM and Glue permissions and a VPC endpoint all in place before an analyst can query S3.
- constraint If a 369 ms planning step recurs on every call, it would outweigh execution for short dashboard queries, and this run cannot show whether it recurs.
The PostgreSQL half of the plan is one node. A Custom Scan sits at the top. Beneath it, Aurora prints a Pushdown SQL line: the query rewritten with kind_id and production_year cast to INTEGER, reading FROM system.main.read_parquet(...) [9]. Under that comes a second engine's operator tree, with READ_PARQUET, PERFECT_HASH_GROUP_BY and count_star() [9]. "Those are DuckDB's names," the post's author wrote [10].
Printing the pushdown text is good engineering. It shows exactly what Aurora handed to the embedded engine, casts included [9]. "I wanted to prove that rather than take it on faith," the author wrote [17]. The available text of the post ends before the pg_duckdb plan appears, so the match itself rests on the post's headline, "Amazon Aurora's analytics is DuckDB" [14].
Getting to that first EXPLAIN takes changes at the platform level. The cluster must run Aurora PostgreSQL 17.11 or later, or 18.6 or later [5]. Its cluster parameter group must set aurora_analytics.enabled = 1 [6]. The IAM role trusts rds.amazonaws.com and needs S3 read on the data bucket, plus Glue read for Iceberg [7]. It attaches through AssociatedRoles with FeatureName: AuroraAnalytics, and the author also added an S3 gateway VPC endpoint [7]. The test ran on a Serverless v2 cluster at 17.11, deployed with CloudFormation [16].
The timings come from one run on one file. Execution took 81.265 ms and planning took 369.237 ms [11]. Planning was about 4.5 times longer [18]. The kind_id = 1 filter ran inside READ_PARQUET. The scan read one file, removed 400,000 rows and returned 100,000 [12], so it kept 20 percent [19].
The cache counters matter most for anyone sizing this. The plan shows 93kB of analytics cache hits, 0kB of remote reads, and zero S3 GET or HEAD requests [13]. For the 81 ms to transfer, your queries would need the same warm cache and files of similar size. I'd expect a cold read to log S3 GETs that this run never made. The post does not explain the planning time. The same plan reports Total CPU Time of 0.0 ms and Effective Parallelism of 0.7 [15], and I would not put that CPU counter on a dashboard yet.
I would run the plan diff on a team's heaviest real query before trusting either the announcement or this post. The reference engine also runs outside AWS: Microsoft put pg_duckdb into public preview on Azure Database for PostgreSQL Flexible Server on November 18, 2025 [1].
What to watch
- The full pg_duckdb plan from the post, and whether its operator tree matches Aurora's line for line on the same query and file.
- Cold-cache runs that log nonzero S3 GET requests, and whether the 369 ms planning time repeats on a second execution.
- Whether AWS documents which DuckDB version aurora_analytics embeds once the DuckLabs acquisition closes.