Skip to content

Build1 publisher3 min readPublished Updated

Fabric's customer-managed keys now reach Spark shuffle and spill, closing a compliance line item

Coverage extending to temporary cluster data removes the exact objection security reviewers used to block Spark workloads. The platform did not change much; the exception request did.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Fabric's customer-managed keys now reach Spark shuffle and spill, closing a compliance line item
Generated illustration

What happened

  • Microsoft Fabric has extended customer-managed key (CMK) encryption to cover temporary data generated during Spark job execution.
  • Fabric's CMK encryption previously covered data at rest: tables, files, and persisted data in OneLake.
  • The gap not covered by CMK was data stored on cluster disks during Spark job execution, shuffle data retained on cluster nodes during processing, and temporary spill files created during large transformations.
  • The GA announcement extends CMK across the entire Spark data lifecycle, from storage through active processing.
  • The post lists five stages of a Fabric Spark job: input read from OneLake (CMK-protected before), shuffle operations between executors (now CMK-protected), spill files written to local disk (now CMK-protected), temporary data cached on nodes (now CMK-protected), and output written back to OneLake (CMK-protected before).

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Microsoft Fabric's customer-managed key encryption now covers the temporary data a Spark job produces while it runs: shuffle exchanged between executors, spill files written to local disk, and data cached on cluster disks [1][3][5]. That is less interesting as cryptography than as paperwork, because those uncovered paths were the specific item security reviewers pointed at when they declined to approve Spark workloads, according to a dev.to write-up of the GA announcement by Gilbert Kiptoo Lelon [15][18].

The old boundary was clean and unhelpful. CMK covered persisted data: tables, files, and anything at rest in OneLake [2]. Everything Spark did between reading input and writing output landed on cluster-local disk under platform-managed encryption instead [6][13]. Spill files in particular are not ephemeral in any way an auditor accepts; the post notes they persist for the duration of the job [7]. Of the five stages the author enumerates in a Fabric Spark job, two were already covered and three now change status [5][1].

The reason this is a blocker story rather than a feature story is the shape of the answer teams had to give. The honest response to "does CMK cover everything" was "everything persisted, but Spark generates temporary data encrypted with platform-managed keys," and that answer triggered risk assessments, legal review, and delay [15]. In banking, where the requirement is often that all customer financial data sit under customer-controlled keys, that gap meant fraud detection, risk modelling and customer analytics on Spark needed a documented security exception [17]. The author's framing is that the change moves Fabric from "requires exception" to "standard approved platform" in many enterprise security frameworks, which is his assessment rather than a vendor commitment [16].

Operationally there is almost nothing to do, which is the right design. The post states that no changes are required to existing Spark jobs or code, that new workspaces get it by enabling CMK at workspace creation, and that workspaces with CMK already enabled pick up the extension automatically [8][9][10]. Encryption happens at the infrastructure layer, transparent to the jobs above it [11]. The key model stays hierarchical: your key in Azure Key Vault, intermediate platform keys derived from it, then data encryption keys on the actual blocks [12]. The consequence worth caring about is the Key Vault audit trail, which the post says now shows key usage across the whole lifecycle [14]. Audit evidence, not the encryption itself, is what closes a control.

Two cautions. This account is a single third-party post, and the reporting carries no GA date, no region or capacity-SKU conditions, and no first-party documentation link [18][2]. And the claim of no performance impact is the author's [10]; encryption on the shuffle and spill path is exactly where a cost would show up, so measure it on your own job shapes before repeating the line in a design review.

What to watch: whether Microsoft's own CMK documentation states the same scope in the same words, including starter pools and any excluded compute; whether the automatic extension to existing CMK workspaces requires a key rotation or workspace restart in practice; and whether your auditors accept Key Vault logs as sufficient evidence for the temporary-data stages, since that is the whole point of the change [14].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories