Build1 distinct publisher3 min readPublished
Mariko, a Principal Applied Scientist at Microsoft, makes recall a gate rather than a tradeable metric. The structure travels well; the thresholds are not published.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
A benchmark score is a comparison instrument, and in prototyping it does real work: ranking models, testing an initial prompt, deciding whether an idea is technically plausible at all [3]. A production decision is a different object. It picks one configuration that maximises a chosen quantity without breaking the things you refuse to break, and no amount of leaderboard movement converts one into the other. Mariko, a Principal Applied Scientist at Microsoft who leads agentic AI work for cybersecurity operations [1], puts the switching point at the approach to production, and names what degrades in between: real inputs are ambiguous, labels may be inconsistent, context may be missing or truncated, and the evaluation set may not reflect the production distribution [4]. Edge cases that barely register in a benchmark become a common source of failure, and improved offline metrics need not translate [5].
What makes the secret-scanning case load-bearing is error asymmetry. Wrongly suppressing a real credential matters more than asking a developer to review one extra alert, so precision and recall were not treated as interchangeable [10]. Recall became a gate instead of a term to be traded away: an experiment could advance only if any recall loss stayed inside a range fixed beforehand [11]. That one structural choice does most of the work, because it means the variant with the best headline precision can be rejected while a weaker one ships [14].
Underneath the gate sit the deployability tests: latency, cost, reliability, production compatibility [12]. They exist to catch the other failure, the change that improves output quality while making the system too slow, too expensive, or too hard to integrate [17]. A team measuring quality alone does not discover that class of rejection until the rollout conversation.
What the post withholds is magnitude. No recall tolerance band, no achieved false-positive reduction, no latency or cost figure for the selected configuration [1]; the tradeoff is demonstrated with two hypothetical experiments instead [14]. That limits what a reader can carry away. The shape of the evaluation transfers, but the numbers cannot, because how much recall you may give up when suppressing credential alerts is a policy call owned by whoever answers for a leaked token [10].
Which is the actual argument behind starting with the product decision rather than the model. The reflex when an LLM system underperforms is to rewrite the prompt, add context, insert another reasoning step, adjust the pipeline, or switch models [9]. Each is cheap to attempt and none of them settles which mistakes you are willing to pay for. Mariko's second practice, treating offline evaluation as integration testing [15], follows from the same position: the thing under test is a system meeting its production inputs, not a model on a clean set [2]. She presents the lessons as portable to code analysis, developer tools, security work, and data analysis [16], and the portable part is the ordering, not the thresholds.
Ranked by verification strength, evidence, and original report placement.
Mariko is a Principal Applied Scientist at Microsoft, where she leads the development of agentic AI workflows for cybersecurity operations.
A language model can perform well on a clean benchmark and still struggle with the cases that matter in production.
Benchmarks and curated datasets are useful when prototyping an LLM-based system: they help teams compare models, test an initial prompt, and determine whether an idea is technically plausible.
As a system moves closer to production the evaluation problem changes: real inputs are often ambiguous, labels may be inconsistent, important context may be missing or truncated, and the evaluation set may not reflect the production distribution.
Edge cases that rarely appear in benchmarks can become common sources of failure, and even when offline metrics improve those results may not translate cleanly into production behavior.
The team encountered these challenges while evaluating an LLM-based system designed to reduce false positives in GitHub secret scanning.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Method fully specified, outcomes unquantified
The evaluation design is described in concrete, checkable detail — criteria tiers, the recall gate, the selection rule, run-metadata capture, one-variable-at-a-time experiments — and comes from the team that ran it. But every quantitative anchor is absent, the one comparison shown is explicitly hypothetical, and there is a single first-party source with no independent corroboration.
One first-party deployment disclosure
There is a real, named production context — LLM-based false-positive reduction in GitHub secret scanning, moved from prototype toward production with a selected configuration — but scope, traffic share, rollout stage, and model identity are undisclosed, and no other organisation is shown adopting the approach.
Slightly overstated: generalisation asserted, numbers withheld
The writing is largely hedged and self-limiting — hypothetical examples are labelled as such, and the framing that benchmarks are prototyping tools is modest rather than promotional. The gap comes from claiming broad cross-domain applicability on the strength of one case study while withholding the recall threshold and the achieved false-positive reduction, which makes the success story unfalsifiable.
Vendor-authored on its own channel about its own product
The post appears on GitHub's own blog, is written by a Microsoft Principal Applied Scientist, and describes an LLM feature inside a GitHub security product. The publisher benefits from portraying its AI evaluation discipline and its secret-scanning alert quality favourably, and it controls which numbers are disclosed — here, none.
Moderate: internally consistent but single-source and unquantified
Confidence in what the post says it did is reasonably high: it is first-party, specific, internally coherent, and candid about which examples are hypothetical. Confidence that the approach produced the claimed production benefit, or that it generalises, is low because there is one vendor source, no metrics, and no independent verification.
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
product
The AI-wrote-it claim died in eight hours. The Actions injection pattern did not.1 distinct publisher
invest
The AI deal frame flipped: buy at 15 times revenue, pay with paper marked at 401 distinct publisher
invest
Cursor ships Origin to paying users as GitHub's outage count reaches 2571 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026