Meta's Prompt Guard 2 caught 6 of 629 buried attacks at a 0.5 cutoff and 621 at 0.003, in a benchmark posted on dev.to. Health checks pass the same at both settings, so only attacks sent through the deployed threshold show which one a team runs.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+30
- Incentives
- Insufficient
- Confidence40
Adversa AI says its Cryptographic Context Injection recovers hostile prompts inside the code sandbox, where filters do not look. xAI has not replied; Google scopes jailbreaks out entirely.
Perspective Coverage
6 publishers
- Builder
- Builder 34%
- Operator
- Operator 55%
- Investor
- Investor 11%
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+30
- Incentives55
- Confidence55
Adversa AI put AES ciphertext on a webpage that no text filter could read. Grok's own code execution decrypted it, and the plaintext came back classified as the model's reasoning rather than as fetched content.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence45
The dev.to post announcing llm-sentinel says llm-guard is archived. The replacement is ten pattern-matching scanners with a score threshold each, benchmarked on 133 hand-written cases its author calls a smoke test.
Reality
- Evidence42
- Adoption8
- Hype gap−12
- Incentives68
- Confidence46
AWS has published a pipeline that pulls Bedrock guardrail traces out of model invocation logs, rewrites them as OCSF Detection Findings, and lands them in the CloudWatch store where analysts already query CloudTrail and VPC Flow Logs.
Reality
- Evidence58
- Adoption22
- Hype gap+12
- Incentives76
- Confidence60
PolicyAware 0.4.4 evaluates identity, tenant, region, arguments and risk tier before an agent's action changes state. Enforcement sits inline on every turn, and the project's own guidance is that adopters measure what that costs.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+18
- Incentives80
- Confidence45
A dev.to writeup moves tool guards out of the prompt and into the executor, where a database check runs before dispatch. The same design fails open when the check itself breaks.
Reality
- Evidence32
- Adoption12
- Hype gap+10
- Incentives78
- Confidence55
With no completed sales, nine live asks running from £35 to £240 and no route to the £26.99 shop price, a self-hosted pricing tool proposed £35.39 and recorded the missing check only in a log line.
Reality
- Evidence50
- Adoption12
- Hype gap+6
- Incentives30
- Confidence55
A dev.to post wraps an autonomous code-review agent in a CLOSED/OPEN/HALF-OPEN state machine. The constructor's defaults show which of its four blast-radius dimensions an adopter has to wire up alone.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+34
- Incentives25
- Confidence45
The F-072 rule engine scans an LLM's chain of thought for risk words and multiplies position size by 0.8 when it finds one. On the trade its author logged, both constants it applied were already in the model's own reasoning.
Reality
- Evidence34
- Adoption8
- Hype gap+45
- Incentives80
- Confidence45
Reconstructing an incident from CloudTrail, AWS Config and Kubernetes events is a correlation problem before it is a language problem, and this engine's timestamp floors decide which links the model may narrate at all.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+12
- Incentives78
- Confidence36
Centralising prompt checks and PII filters at the gateway mirrors putting TLS at the edge, and the arithmetic backs that comparison. The semantic checks themselves, though, work nothing like TLS termination.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+34
- Incentives76
- Confidence36
TraceSafe-Bench ran 20 guard systems over more than a thousand edited tool-call traces, and the ranking followed structured-data ability rather than safety alignment. Your prompt-injection number does not cover the runtime.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+26
- Incentives55
- Confidence38
Model-level guardrails check the prompt and the response, which leaves the parameters the model just chose for a tool unexamined. AWS's answer is three lifecycle hooks, and the per-turn cost is yours to measure.
Reality
- Evidence56
- Adoption15
- Hype gap+18
- Incentives82
- Confidence54
ChatGPT for Teens folds a year of bolt-on parental controls into a product boundary. The hard part is now deciding which account a minor is sitting in, and who configured it.
Reality
- Evidence34
- Adoption30
- Hype gap+18
- Incentives72
- Confidence44
Adversa says it hid data-exfiltration instructions in AES-256-GCM ciphertext and let Grok decrypt them in its own Python sandbox. The plaintext version of the same attack was refused.
Reality
- Evidence42
- Adoption20
- Hype gap+22
- Incentives68
- Confidence44