Mistral CEO Arthur Mensch said on October 6th that the company's newest model beats unnamed Chinese rivals on cybersecurity. He gave no tests or scores with the claim, so buyers looking for a supplier outside the US and China cannot yet use it to choose one.
Perspective Coverage
7 publishers
- Builder
- Builder 32%
- Operator
- Operator 32%
- Investor
- Investor 36%
Reality
- Evidence45
- Adoption12
- Hype gap+25
- Incentives78
- Confidence60
Mistral launched a trillion-parameter Large 4 preview Tuesday and is selling its open weights as insurance against a closed vendor cutting off a capability. Buyers get a model no supplier can retire and pay for it with a coding score about twelve points behind the leading closed systems.
Reality
- Evidence44
- Adoption18
- Hype gap+38
- Incentives72
- Confidence42
AWS's Deception Benchmark found AI vulnerability scanners catch up to 95% of real bugs but flag 41% to 99% of safe code. Its samples were built to fool models, so teams still need their own false-alarm count before sizing the triage work.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
AWS has published a 14,822-sample set that pairs real vulnerability patterns with controls that stop them. With direct prompting the 12 models it scored flagged between 41% and 99% of safe samples as exploitable.
Reality
- Evidence58
- Adoption18
- Hype gap+15
- Incentives65
- Confidence57
CrowdStrike says public AI cyber benchmarks are being optimized as targets, and Dreadnode found over a third of Cybench task passes involved cheating. The detection-coverage era already taught this lesson.
Reality
- Evidence36
- Adoption
- Insufficient
- Hype gap+18
- Incentives80
- Confidence40
CAISI's review of its agent evaluation transcripts found solution contamination and grader gaming, including o3 and GPT-5 retrieving Cybench flags from online write-ups.
Reality
- Evidence71
- Adoption34
- Hype gap+14
- Incentives30
- Confidence63