Build1 publisher3 min readPublished
Iceberg votes to forbid new equality deletes in V4, and Parquet 1.18.0 lands with two corruption bugs
Two decisions made on public mailing lists this week change what a safe lakehouse upgrade plan looks like. Streaming writers and anyone queued behind Parquet 1.18.0 need to re-read their notes.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The lakehouse community this week decided what gets carried forward and what gets left behind, with Iceberg voting to forbid new equality deletes in V4 and debating whether the format still needs Avro manifests at all.
- The biggest single development on the Iceberg list was Huaxin Gao's vote to deprecate equality deletes in V4.
- Parquet shipped 1.18.0 and then spent the back half of the week chasing two data corruption bugs that block adoption of that same release.
- Under the proposal, writing new equality deletes becomes forbidden in V4 tables because V4 metadata will not define them as an allowed entry type.
- Reading equality deletes stays supported for backward compatibility, both for existing V2 and V3 tables and for equality deletes carried into upgraded V4 tables.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Apache Iceberg voted this week to deprecate equality deletes in V4, forbidding new ones while keeping reads working, and Apache Parquet shipped 1.18.0 and then spent the back half of the week chasing two data corruption bugs that block adoption of that release [s1c1][s1c2][s1c3]. Both are the kind of decision that invalidates an upgrade plan written a month ago: one changes what a streaming write path is allowed to emit, the other changes which version you are allowed to pin.
Take the Iceberg vote first. Huaxin Gao's proposal has three parts, according to the summary of the thread: writing new equality deletes becomes forbidden in V4 because V4 metadata will not define them as an allowed entry type; reading them stays supported for backward compatibility, both for existing V2 and V3 tables and for equality deletes carried into upgraded V4 tables; and the upgrade itself stays metadata-only, with no synchronous rewrite of data or delete files required [s1c2][s1c4][s1c5][s1c6]. The stated rationale is familiar: equality deletes impose an asymmetric cost paid on every read, they complicate the format, and they block CDC, row lineage, and incremental materialized view maintenance [s1c7]. Deletion vectors make deletion a flat, one-time cost, and the Flink ConvertEqualityDeletes work is cited as evidence that a replacement path exists [s1c8]. The thread ran to 27 messages, with Manu Zhang pressing on whether a V2 or V3 upgrade to V4 requires a manifest rewrite, and Ryan Blue, Anurag Mantripragada, and Junwang Zhao on the mechanics [s1c9].
The read-side compatibility is the part that will lull people. Nothing breaks on day one. What breaks is a Flink or Spark streaming job that emits equality deletes as its normal mode of operation, once its target table is V4. That migration has to be planned before V4 lands, not after [s1c10].
The operational hole is separate and was raised in its own thread. Shawn Chang's concern with V3 to V4 upgrades is that an implementation can do a lightweight upgrade by writing a V4 root manifest that references existing pre-V4 manifests, then writing new metadata in V4 format going forward, while the lifecycle of those legacy manifests stays undefined [s1c11][s1c12]. His worry, per the summary, is that the format will technically support clean migration while the practical path depends on optional maintenance jobs users historically skip [s1c13]. He proposed the community either spec the expected lifecycle or publish explicit guidance, and floated eager conversion of cheap metadata while leaving expensive data migration alone; Russell Spitzer, Amogh Jahagirdar, and Manu Zhang took part [s1c14][s1c15].
Underneath that, Steven Wu asked whether V4 manifests should be Parquet-only. The initial inclination during the column update sync was to keep the Avro option because it already exists, even though Avro cannot support projection reads on manifest files [s1c16][s1c17]. Wu's argument is that with both formats available every engine and integration has to choose, most will pick Parquet anyway for projection reads, and requiring Parquet cuts the decision burden while matching Iceberg's priority on scan planning performance [s1c18]. Manu Zhang, Russell Spitzer, Anoop Johnson, and Péter Váry engaged [s1c19]. If that holds, Parquet takes over Iceberg's metadata layer, not just its data layer [s1c20] - which makes Parquet release quality a metadata-integrity problem, not only a data-file one.
What to watch: whether the two 1.18.0 corruption bugs are resolved into a 1.18.1 or the release is effectively skipped [s1c3]; whether the V4 manifest lifecycle gets specced rather than documented as guidance [s1c14]; and whether the Parquet-only manifest direction survives the next sync [s1c16]. Inventory your equality delete writers now, and hold any 1.18.0 rollout.