Build1 publisherNot yet confirmed elsewhere2 min readPublished
Salesforce Data 360 injects a Spark listener plugin at its job gateway to find idle executors
Salesforce Data 360 instruments about eight million Spark runs a day at one submission gateway, without changing application code. The plugin fails open, so a broken collector leaves the platform team a data gap to chase while customer jobs keep running.
The Engineer · Build desk

What happened
- One Data 360 workload requested 40 executors at startup while most had no useful work, producing pod churn, node churn and idle capacity with no job failure.
- Fixed resource profiles and tenant-and-data-volume estimates both set an allocation before submission, and neither showed how that allocation behaved once the job ran.
- Listeners in each Spark driver collect the telemetry and export it through Kafka into a central Iceberg lakehouse.
- Alerts watch for export failures and telemetry gaps so the platform team can investigate evidence that went missing.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The approach needs one submission path that every Spark job crosses; a platform where jobs arrive through several entry points has no single place to add the plugin and is back to asking each team.
- decision Teams that size Spark jobs from tenant or data-volume estimates now face a choice: build execution-time measurement first, or cut executors on a guess.
- exposure A sizing query fed by incomplete applications, missing records or miscounted retries can point at waste that is not there, so the telemetry needs its own checks before it drives a cut.
The four authors state the premise in one sentence. They wrote, "Successful completion is evidence of a working configuration, not an efficient one." [1][7] Their workloads include short runs and also jobs that process terabytes over several hours [2]. Averaged across a day, eight million runs is about 93 Spark applications every second [18].
The evidence they wanted was per executor. Which executors received tasks, how much CPU and memory each used, and whether execution spilled or retried [8]. Those answers let the team separate unused capacity from memory pressure and shuffle bottlenecks [3]. If you cut without them, the saving can turn into slower runs, expensive retries or production failures [4].
Application teams own the code. Asking each of them to add instrumentation creates an integration project before anyone can look at the waste [9]. Data 360 already had a single entry point. The Data Processing Controller, or DPC, is the common gateway for managed Spark submissions behind ingestion, segmentation and activation [10]. Spark's listener hooks observe execution independently of application code, so the team extended the open-source DataFlint plugin and had DPC inject it as it built each submission [11].
The listeners run on a dedicated listener-bus queue [13]. If the instrumentation cannot be configured or initialised, the submission proceeds anyway [13]. I think fail-open is the right default for a platform team that does not own the applications. The authors set out to avoid giving those applications another way to fail [9], and a collector able to block a submission would be one.
The authors asked whether the 40-executor job needed that capacity immediately, or could start smaller and add executors as demand appeared [17]. Their sections on collection and interpretation do not include the answer, or a savings figure.
What to watch
- Whether Salesforce publishes the executor, CPU and memory reductions, and the retry rates, that followed from acting on the telemetry.
- Whether the 40-executor workload moves to a smaller starting allocation that adds executors as demand appears.
- Whether the team's DataFlint extensions are contributed to the open-source project for other Spark platforms to reuse.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence38
- Adoption30
- Hype gap+5
- Incentives40
- Confidence45
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The post was written by Siddharth Sharma, Suvhrajit Basak, Kusum Dhalia, and Nabeel Qaiser.
- [2]
Salesforce Data 360 handles roughly eight million Apache Spark workload executions a day, ranging from short runs to workloads processing terabytes over several hours.
- [3]
The team needed to distinguish unused capacity from memory pressure and shuffle bottlenecks, preserve performance and reliability, and collect the evidence without requiring application code changes.
- [4]
Successful execution does not tell you how much allocated capacity a workload needed; cutting CPU, memory, or executors without that evidence can turn potential savings into slower runs, expensive retries, or production failures.
- [5]
A Data 360 workload's configuration requested 40 executors at startup, yet most had no useful work initially; such an allocation creates pod churn, node churn, and unused capacity without producing a job failure.
- [6]
Neither a fixed resource profile nor an estimate based on tenant and data volume, an approach other workload owners used, could answer whether the job needed that capacity; both helped choose an allocation before submission but neither revealed how the allocation behaved during execution.
- [7]
"Successful completion is evidence of a working configuration, not an efficient one."
- [8]
For the 40-executor workload the team needed to know which executors received tasks, how much CPU and memory they used, and whether execution spilled or retried.
- [9]
Application teams own the code; asking every team to add instrumentation creates an integration project before the waste can be investigated, and the team wanted evidence without introducing another way for applications to fail.
- [10]
The Data Processing Controller (DPC) is the common gateway for managed Spark submissions supporting Data 360 ingestion, segmentation, and activation.
- [11]
Spark listener hooks let the team observe execution independently of application code; the team extended an open-source DataFlint Spark plugin and injected it when DPC constructed a submission.
- [12]
The listeners collected telemetry in each Spark driver and exported it through Kafka into a central Iceberg lakehouse.
- [13]
Listeners used a dedicated listener-bus queue, and submissions continued if instrumentation could not be configured or initialized.
- [14]
Alerts monitored export failures and telemetry gaps so the platform team could investigate missing evidence.
- [15]
The authors advise readers to look for an equivalent shared boundary in their own platform to provide consistent Spark observability while keeping collection isolated from customer computation.
- [16]
Incomplete applications, missing records, or miscounted retries can distort an apparent optimization candidate before any setting is changed.
- [17]
The team asked whether the 40-executor application needed that capacity immediately or could start smaller and add executors as demand appeared.
- [18]
Eight million Spark runs a day averages about 93 runs per second.
Sources
1 independent publisher whose own reporting we read for this story.
- engineering.salesforce.comApache Spark Resource Optimization: Lessons From Eight Million Jobs a Day
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.