Measured five times on the same plan, one safety critic returned five different verdicts and still never approved a seeded defect. It holds because the code downstream never reads the model's severity label.
Reality
- Evidence40
- Adoption12
- Hype gap+8
- Incentives45
- Confidence52
The AI Frontier Division built Cofa-Probe on OpenAI's Codex to score four stages of human-agent collaboration, which tells you roughly what Krafton now thinks a finished portfolio piece is worth as hiring evidence.
Reality
- Evidence36
- Adoption31
- Hype gap+18
- Incentives71
- Confidence44
A developer ran 200 flagged snippets past two frontier models with identical prompts. One cleared 51% of the false alarms; the other cleared 20% and agreed with 90% of what it saw. The countermeasures are a model property.
Reality
- Evidence24
- Adoption9
- Hype gap+32
- Incentives46
- Confidence33
A paper on arxiv.org names the failure the Echo Gap: wrong episodes get inflated scores, then get retrieved more. The proposed fix needs a verifier whose errors are uncorrelated with the first grader's.
Reality
- Evidence34
- Adoption12
- Hype gap+22
- Incentives62
- Confidence44
Five models, ten questions, a tidy leaderboard. Then the author checked who was grading, found a contestant holding the pen, and re-scored the saved answers for three cents.
Reality
- Evidence55
- Adoption15
- Hype gap−10
- Incentives30
- Confidence52