Skip to content

Build1 publisherNot yet confirmed elsewhere2 min readPublished

Salesforce Data 360 injects a Spark listener plugin at its job gateway to find idle executors

Salesforce Data 360 instruments about eight million Spark runs a day at one submission gateway, without changing application code. The plugin fails open, so a broken collector leaves the platform team a data gap to chase while customer jobs keep running.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Salesforce Data 360 injects a Spark listener plugin at its job gateway to find idle executors
Generated illustration

What happened

  • One Data 360 workload requested 40 executors at startup while most had no useful work, producing pod churn, node churn and idle capacity with no job failure.
  • Fixed resource profiles and tenant-and-data-volume estimates both set an allocation before submission, and neither showed how that allocation behaved once the job ran.
  • Listeners in each Spark driver collect the telemetry and export it through Kafka into a central Iceberg lakehouse.
  • Alerts watch for export failures and telemetry gaps so the platform team can investigate evidence that went missing.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The approach needs one submission path that every Spark job crosses; a platform where jobs arrive through several entry points has no single place to add the plugin and is back to asking each team.
  • decision Teams that size Spark jobs from tenant or data-volume estimates now face a choice: build execution-time measurement first, or cut executors on a guess.
  • exposure A sizing query fed by incomplete applications, missing records or miscounted retries can point at waste that is not there, so the telemetry needs its own checks before it drives a cut.

The four authors state the premise in one sentence. They wrote, "Successful completion is evidence of a working configuration, not an efficient one." [1][7] Their workloads include short runs and also jobs that process terabytes over several hours [2]. Averaged across a day, eight million runs is about 93 Spark applications every second [18].

The evidence they wanted was per executor. Which executors received tasks, how much CPU and memory each used, and whether execution spilled or retried [8]. Those answers let the team separate unused capacity from memory pressure and shuffle bottlenecks [3]. If you cut without them, the saving can turn into slower runs, expensive retries or production failures [4].

Application teams own the code. Asking each of them to add instrumentation creates an integration project before anyone can look at the waste [9]. Data 360 already had a single entry point. The Data Processing Controller, or DPC, is the common gateway for managed Spark submissions behind ingestion, segmentation and activation [10]. Spark's listener hooks observe execution independently of application code, so the team extended the open-source DataFlint plugin and had DPC inject it as it built each submission [11].

The listeners run on a dedicated listener-bus queue [13]. If the instrumentation cannot be configured or initialised, the submission proceeds anyway [13]. I think fail-open is the right default for a platform team that does not own the applications. The authors set out to avoid giving those applications another way to fail [9], and a collector able to block a submission would be one.

The authors asked whether the 40-executor job needed that capacity immediately, or could start smaller and add executors as demand appeared [17]. Their sections on collection and interpretation do not include the answer, or a savings figure.

What to watch

  • Whether Salesforce publishes the executor, CPU and memory reductions, and the retry rates, that followed from acting on the telemetry.
  • Whether the 40-executor workload moves to a smaller starting allocation that adds executors as demand appears.
  • Whether the team's DataFlint extensions are contributed to the open-source project for other Spark platforms to reuse.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence38
Adoption30
Hype gap+5
Incentives40
Confidence45
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The post was written by Siddharth Sharma, Suvhrajit Basak, Kusum Dhalia, and Nabeel Qaiser.

    ReportedSupportedSource: Salesforce Engineering blog bylineView cited source
  2. [2]

    Salesforce Data 360 handles roughly eight million Apache Spark workload executions a day, ranging from short runs to workloads processing terabytes over several hours.

    ReportedSupportedSource: Salesforce Engineering blogView cited source
  3. [3]

    The team needed to distinguish unused capacity from memory pressure and shuffle bottlenecks, preserve performance and reliability, and collect the evidence without requiring application code changes.

    ReportedSupportedSource: Salesforce Engineering blogView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. engineering.salesforce.com

    1 article · October 8, 2026

    Apache Spark Resource Optimization: Lessons From Eight Million Jobs a Day

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Spark resource right-sizingFollow
  • Data platform observabilityFollow
Loading related stories