Build1 distinct publisher2 min readUpdated
Microsoft archived its prompt-injection framework on March 27, 2026. The tools that remain split into app-layer and model-layer work, and PyRIT was the only one configurable for both.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
An archived repository keeps working. Nothing uninstalls itself the day the maintainer stops answering issues, which is why this sort of change spreads slowly and badly: the package still runs and the docs still read as current, while the issue tracker has quietly stopped receiving anything [2]. The dev.to comparison carrying the notice is one of the pages that got corrected, and even there the update is dated roughly four months after the archive [16]. PyRIT's audience was security researchers running structured red-team engagements rather than app developers doing a pre-ship check, and the guide puts its general-developer adoption noticeably below the other four [14]. There was no large user base to notice out loud.
Treating the survivors as substitutes is the second error. garak is pip-installable, ships 50+ probes, carries 8.1k GitHub stars and is described as actively maintained by NVIDIA [9], and none of that helps if what you ship is an application, because a model-layer scan never touches your system prompt, your tool permissions or your retrieval path [11]. promptfoo runs the other way: 50+ red-team plugins, zero install via npx, and report presets already mapped to OWASP LLM Top 10, NIST and MITRE ATLAS [6]. The same guide is honest that the plugin surface is large enough that a useful first run takes well past five minutes [8]. Giskard has OWASP-mapped detectors as well, but the continuous-scan Hub, the part that would let you run it repeatedly against a live app, is a paid product, leaving the free tier as a point-in-time check [12]. sentinel-scan-cli is placed on the app-layer side alongside those two [5].
Two notes on the evidence, since all of this comes from one author. The claim that promptfoo is used internally at OpenAI and Anthropic is attributed to promptfoo's own repository [7], which is a vendor self-report and should be read as one. And the comparison closes by routing readers to its own separate post on what to use instead [18]. Neither weakens the layer taxonomy, which is the most useful thing in the piece: decide whether you are testing the wrapped application or the raw model, and most of the tool choice is already made [4]. The place to distrust a fast answer is the multi-turn gap. PyRIT was a framework for scripting escalating conversational attacks rather than a turnkey scanner [13], and the guide's own suggestion is that promptfoo's red-team plugins cover a lot of the same ground [15]. A lot is not all, and the difference is the part nobody has rebuilt.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The guide states promptfoo is used internally at places like OpenAI and Anthropic, attributing that to promptfoo's own repository.
Microsoft archived PyRIT on GitHub on March 27, 2026.
The PyRIT repository is now read-only: no commits, no releases, no issue triage, and whatever version a user has installed is the last one they will get.
The guide added an update dated August 2026 recording the archive, and kept its PyRIT section because a lot of existing guides and tutorials still point people to the tool.
The guide distinguishes app-layer testing (whether the application's prompts, guardrails, tool-calling logic and RAG retrieval resist attacks wired together end to end) from model-layer testing (whether the underlying model has exploitable behavior independent of any app around it).
promptfoo, Giskard and sentinel-scan-cli are primarily app-layer tools; garak is model-layer; PyRIT could do either depending on how it was configured.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single practitioner source, specific but uncorroborated
Every claim traces to one self-published dev.to guide by an author who discloses working on one of the five tools compared. The load-bearing facts are unusually specific and checkable (archive date, star count, probe/plugin counts, paid-tier boundary) and the guide states its method (ran tools where possible, read source/docs otherwise), which lifts it above assertion. But there is no second publisher, no primary GitHub or vendor artifact in the cluster, and the strongest adoption claim is explicitly vendor-attributed — so the evidence base is credible-but-thin.
Concrete signals for the survivors, negative signal for PyRIT
There are real adoption datapoints: garak at 8.1k stars with active NVIDIA maintenance, promptfoo described as the most broadly adopted with vendor-reported internal use at OpenAI and Anthropic, and PyRIT explicitly characterized as having noticeably lower general-developer uptake before it was archived. What is missing is any figure for promptfoo, Giskard or sentinel-scan-cli comparable to garak's star count, and no data on how many teams actually have PyRIT in a pipeline today — which is the number that would size the migration problem.
Mildly overstated framing over sound underlying facts
The claims themselves are hedged and mostly concrete, and the guide discloses the author's conflict of interest, which pulls the gap toward zero. It tips slightly positive because the framing outruns the evidence in two places: 'used internally at OpenAI and Anthropic, which tells you it holds up at scale' is a vendor-repo assertion doing inferential work, and the headline urgency of an archive is tempered by the guide's own admission that PyRIT had the lowest general-developer adoption of the five. The delayed update note — roughly four months after the archive — also argues the disruption was less acute in practice than the framing implies.
Disclosed vendor stake in the comparison set
The author states outright that they work on sentinel-scan-cli, one of the five tools being ranked, and the guide's structure gives that tool a distinct 'speed to first result' niche while conceding it covers 15 attack patterns rather than 50+. Disclosure plus stated limitations materially mitigate the conflict, but the incentive is real and structural: the comparison's author benefits from how the field is partitioned. Secondary commercial interests appear in Giskard's paid continuous-scan Hub and in promptfoo's self-reported adoption claim.
Moderate on the archive fact, weak on the surrounding comparison
Confidence is asymmetric. The archive of PyRIT is dated, specific, repeated twice in the source and consistent with the guide's editorial behaviour, so it is likely accurate. The comparative judgements — layer assignments, relative adoption, setup-time claims, absence of a one-for-one replacement — rest on one practitioner's testing and reading, filtered through a disclosed stake in one of the tools, with no independent corroboration and no benchmark results in the cluster.
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
invest
The 81% Problem: AI's Star CEOs Are Polling Badly With The People They Need To Hire1 distinct publisher
product
AMD borrows $4.75bn while sitting on $13bn, and the number matches its Anthropic promise1 distinct publisher
invest
The AI moat is now a balance sheet, so price the financing and not the model1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 23, 2026