Skip to content

Build1 publisher3 min readPublished

Iceberg's REST catalog gains a nine-action vocabulary for handing enforcement to trusted engines

Apache Iceberg's REST Catalog spec now lets a catalog return per-user row filters and nine column actions for a trusted engine to enforce, Databricks says. Policies that need subqueries or lookup tables do not fit that vocabulary, so the portability covers simple rules running on engines a deployment already trusts.

The Engineer · Build desk

Illustration accompanying Iceberg's REST catalog gains a nine-action vocabulary for handing enforcement to trusted engines

What happened

  • Catalog labels, the second addition, are meant to make governance context portable from one catalog to another.
  • Under read restrictions the engine gets the outcome of a policy for one principal, as filtering or masking instructions, never the policy the administrator wrote.
  • Databricks treats Spark and DuckDB as untrusted when users control the runtime, because those users can run arbitrary code or read the data directly.
  • The spec defines what a trusted engine must enforce but leaves out how a catalog establishes that trust, and a client's own claim is not enough.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Rules built on subqueries, lookup tables or custom UDFs lose their meaning when reduced to the standard vocabulary, so only simple policies can be handed to an outside engine intact.
  • decision Teams whose users run their own Spark or DuckDB cannot delegate enforcement to those engines and must keep catalog-side filtering in the read path for governed data.
  • exposure Each deployment owns identity propagation and credential binding, so a weak trust setup leaves restrictions with an engine that can ignore them.

A governed read under the new contract starts when a reader loads a table. The catalog evaluates the policies that apply to the requesting principal and the request context. It returns required column-projection actions and row-filter expressions, and the trusted engine must apply them as it reads [7]. The initial vocabulary is nine predefined column actions plus standardized row filters such as comparisons and set membership [9]. I think that is the right size for a first version. An engine team can implement nine actions and a small predicate language and test every case.

The hard part is reduction. The catalog has to reduce each policy to that vocabulary before it answers, and many enterprise policies depend on subqueries, lookup tables or custom UDFs that the vocabulary cannot express [10]. Databricks wrote that "policies can be represented only when the catalog can reduce their result to the vocabulary defined by the standard, otherwise you lose policy semantics" [10].

Delegation also depends on the catalog being able to rely on the engine to enforce the restrictions and stop users from bypassing them [4]. A securely configured Trino deployment is Databricks' example of an engine that qualifies, because it enforces row filters and column masks natively [6]. A user who can run arbitrary code against the files will not be slowed down by a list of restrictions in a load-table response [5]. Community discussion has considered mTLS and OAuth for establishing trust, but trust remains outside the protocol [12]. Databricks lists three open implementation questions. One is how an engine securely propagates the end user's identity and attributes. The others are how a catalog tells the user apart from the engine acting for that user, and how credentials are bound to their intended recipient [14].

Databricks also describes the path that avoids delegation. On dedicated compute it routes queries through a secure filtering fleet [15]. Unity Catalog's cross-engine ABAC puts that fleet behind the Iceberg REST scan/plan APIs, so data is sanitized before an external engine such as Spark or DuckDB processes it [3]. The post recommends read restrictions for direct engine-to-catalog access where the source catalog's policies are simple and a trusted engine can enforce the result [13]. Enforcement that works across vendors needs both conditions to hold for a team. Its policies must reduce to comparisons, set membership and the nine column actions. Its engines must run where the catalog operator can trust them. The account of where that line falls comes from Databricks, whose Unity Catalog offers the centralized alternative [3].

Catalog labels are the second addition the Iceberg community advanced [1], and they get far less detail in the Databricks account available here. It does not describe the label format or what a receiving catalog must do with a label, so the claim that labels carry governance between catalogs cannot yet be checked against spec text.

What to watch

  • Whether the Iceberg spec widens the read-restrictions vocabulary past nine column actions and simple row filters to cover subquery or lookup-table policies.
  • Whether the community standardizes engine trust and identity propagation, now discussed via mTLS and OAuth but kept outside the protocol.
  • Publication of the catalog-labels format and which engines beyond Trino implement read restrictions.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories