Build1 distinct publisher2 min readPublished
Its engineers describe 12,000 dashboards, 20-plus apps and an accurate churn score no one could act on. The fix was a decision layer, and only one of its three inputs comes from a model.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Read the three inputs closely and one of them is doing less work than its billing suggests. Of signal, business logic and knowledge, only the signal comes out of a model [7]. Thresholds, priorities, operating rules and current business motions are things a revenue organisation decides and keeps rewriting. Institutional and field expertise about what a good response looks like sits with people who were never in the training pipeline [7]. Two of the three ingredients in the Next Best Action layer have owners outside the data science team [12].
That is what makes the 0.83 example useful rather than cute. The churn model was not wrong; by the authors' account it did its job, and the open questions were whether that number was high enough to act on, why the account was at risk, whether it outranked another renewal, and what move to make [4]. The team had years of XGBoost and gradient boosting work behind them on seller performance, pipeline health, skill gaps, attrition risk, product adoption and onboarding completion, all good at assessment and stopping there [5]. Retraining does not produce a priority order. Somebody in sales operations has to write one down.
The dashboard count deserves arithmetic the post does not do. Roughly 12,000 dashboards [2] at one minute each, without pauses, is 200 hours, or 25 eight-hour days of looking [13]. No consolidation programme survives contact with that number, which is presumably why consolidating them was listed as an option and then set aside [11].
Note what the fix does to the inventory. The stated destination is a single conversation the seller already has open [9]. Nothing in the writeup says the 12,000 dashboards were retired. The delivery surface changed; the estate behind it stays, which means two versions of the truth to keep aligned, and the newer one now asserts an expected impact it can be held to [8].
The supplied text ends mid-sentence, in the middle of a list of diagnostic questions, and carries no adoption figure and no measured outcome [14]. So read it as a design argument rather than a result. The part that transfers is cheap and awkward: take any prediction in production and ask what decision a user can make from that output alone [6]. Applied honestly, the answer is usually that somebody downstream has been supplying the missing interpretation for free, and their time never appeared in the model's cost line. Salesforce's engineers call the reusable pattern the separation of assessment from action, not the layer they happened to build [10].
Ranked by verification strength, evidence, and original report placement.
Closing the gap required three inputs: the signal (model output across performance predictions, pipeline health, skill gaps, attrition risk, product adoption, onboarding completion); business logic (thresholds, priorities, operating rules and current business motions); and knowledge (institutional and field expertise about what good actually looks like in context).
Combining the three inputs allowed the Next Best Action layer to produce a recommendation containing the action, the account it applies to, the reasoning behind it, and the expected impact.
The post 'From Prediction to Action: How to Turn AI Outputs Into Decisions' was published on engineering.salesforce.com by Ali Nahvi, Akshit Behera, Yvonne Fan, and Melissa Ramey.
The problem Salesforce faced in early 2025: sellers had roughly 12,000 dashboards competing for their attention alongside more than 20 additional apps and tools.
The authors state that nothing was technically broken - models accurate, pipelines healthy, dashboards working - yet the people consuming those signals still had to determine which ones mattered, what they meant, and what to do next.
The post uses a churn model returning a risk score of 0.83: the model has done its job, but the user does not learn whether 0.83 is high enough to act on, why the account is at risk, whether to prioritise it over another renewal, or what move to make.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single first-party account, specific figures, no measurement
Everything rests on one self-published post by the team that built the system. It is unusually specific about the problem surface (roughly 12,000 dashboards, 20+ apps, the 0.83 churn example, named model families) and about the architecture's three inputs, which makes the reasoning inspectable. But there is no third-party corroboration, no evaluation methodology, no before/after measurement, and the supplied text terminates mid-sentence, so the central assertion that the decision layer fixed the failure is unverified.
No adoption figures disclosed
The only adoption-adjacent material is a qualitative first-party statement that the layer exists and serves sellers. No user counts, rollout percentage, retirement of the prior dashboards, usage duration, or outcome metric is supplied, and the cluster contains no external deployment, benchmark or pricing evidence. Adoption cannot be scored without inventing facts.
Modestly overstated: strong framing, zero measurement
The piece is more disciplined than most vendor engineering writing - it explicitly says the reusable pattern is the separation of assessment from action rather than its own layer, and it names knowledge extraction as an unsolved, expensive bottleneck. The overstatement is narrower: the transformation from 12,000 dashboards to a single conversation is asserted as an accomplished improvement while no outcome, adoption or comparison figure supports it, and the framing that 'the model was never the problem' is presented as a general lesson from an unmeasured single case.
Vendor-authored, self-assessed, no counterweight
The single publisher is the subject: Salesforce engineers describing a Salesforce internal system on Salesforce's engineering blog, where such posts serve recruiting and credibility for the company's AI positioning. The authors both diagnose the failure and grade their own fix, with no external reviewer, no dissenting source and no disclosed metric that could falsify the account. The incentive is partially offset by the post carrying no product pitch, pricing or upsell language and by its admission that knowledge extraction remains costly.
Confident about what was said, not about what it achieved
Confidence is high that the post says what the ledger records - the figures, the three inputs, the signal-versus-answer test and the stated reusable pattern are all quoted directly and unambiguously. Confidence is low that the described layer produced the implied improvement, because that rests on one unverified, self-interested, unmeasured and textually truncated source with no adoption data.
invest
Salesforce's double digits, minus Informatica: agentic AI is real and still 2% of revenue1 distinct publisher
build
Notion's agent stack is live, not slideware, and it only changes one of your decisions1 distinct publisher
build
A green build only proves your agent was consistent with itself1 distinct publisher
invest
Two judges, 42 hours: Nvidia's print and Warsh's first keynote price the same trade1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 24, 2026