Skip to content

Build1 publisher2 min readPublished

S3 Tables adds geometry, geography and nanosecond timestamps to complete its Iceberg V3 support

Amazon S3 Tables now supports every Apache Iceberg V3 data type after adding four, including geometry, geography and nanosecond timestamps. Telemetry and mapping teams can stop storing event times and coordinates as integers and strings.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying S3 Tables adds geometry, geography and nanosecond timestamps to complete its Iceberg V3 support
Generated illustration

What happened

  • Existing V2 tables can be upgraded to V3 in place, and S3 Tables keeps running compaction and maintenance on them.
  • AWS's What's New post says deletion vectors, row lineage and the variant type were already supported before this release.
  • Column default values also arrive, filling a newly added column for existing rows with no backfill.
  • The new types are available in every AWS Region where S3 Tables is offered.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction AWS's blog lists deletion vectors and variant under 'starting today', but its What's New post calls them existing support. The new reasons to upgrade now are the geospatial types, nanosecond time and column defaults.
  • cost Ingest jobs pay for the cheaper JSON reads. Shredding and statistics collection happen at write time, before any query benefits.
  • capability Incremental pipelines can pick up changed rows from lineage fields the table maintains itself, with no full-table scan.

AWS leads with a compliance delete. Under V2, according to the AWS blog, removing 50,000 user records from a 2-billion-row table leaves positional delete files behind, and queries stay slow until compaction runs [6]. Those records are 0.0025% of the table [1]. V3 swaps positional delete files for deletion vectors, a compact binary format [7]. AWS says the same delete then writes one deletion vector file instead of thousands of small ones [7].

AWS wrote that the change means "significantly reducing compaction time and delete file overhead" [8]. The post does not include compaction times, query latencies or an exact file count [7]. For the example to hold on another table, that team's V2 deletes would have to fan out into thousands of positional delete files today. The delete-file count left after a team's own last compliance run shows whether they do.

Variant targets the JSON-as-string pattern, where every V2 query has to parse each event [17]. During writes, the engine shreds variant values into hidden columns and collects statistics on them. At query time, those statistics let it prune files [9]. The example table sets its version with one property, `TBLPROPERTIES ('format-version' = '3')`, and declares `payload variant` [11]. Inserts go through `PARSE_JSON`. Reads in Amazon EMR Spark use `variant_get` with no parse at read time, and the sample query filters on the payload's `action` field [11]. AWS wrote that pruning "significantly reduces I/O compared to parsing JSON strings" [10]. That claim transfers to workloads whose queries filter on shredded fields, as the example's does.

I think row lineage is the best-engineered piece of the set. The change marker is part of the format: every record gets `_row_id` and `_last_updated_sequence_number` automatically [12]. No pipeline has to stamp them.

The four types that are new in this release replace data V2 tables kept as strings or integers [3] [2] [17]. Geometry and geography hold points, lines and polygons, so fleet-tracking and asset-mapping workloads can filter on location at query time [13]. Unknown, true to its name, is for columns with no known type [18].

What to watch

  • Published compaction and query timings comparing deletion vectors with V2 positional deletes on S3 Tables, from AWS or independent users.
  • Which engines besides Amazon EMR Spark read and write variant, geometry and nanosecond columns on upgraded S3 Tables.
  • Whether AWS documents a path back to V2 for a table upgraded in place.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories