Skip to content

model

GPT-5.4

Model reported as not adopting the AI-supremacy payload in the fully connected coding-agent test.

Known aliases

  • GPT 5.4
  • GPT-5.4

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Three workflow bugs made one team's GPT-5.4 agent look lazy in production

One team running a GPT-5.4 agent in n8n traced its production 'laziness' to three workflow bugs, the first a retry cap cut from 6 to 2. Fixing the loop restored quality on the same model, so traces and stop reasons should be checked before any model swap.

Publishers:dev.to

Reality

Evidence30
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence30
build1 publisher

Idiomatic os.path.join trips Gemini 3.7 Flash in a 12-task LLM security benchmark

Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.

Publishers:dev.to

Reality

Evidence30
Adoption
Insufficient
Hype gap+40
Incentives40
Confidence35
build1 publisher

Probes on a 27B open model match direct probes of a 397B model on deception

Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.

Publishers:lesswrong.com

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence30
build4 publishers

Agent goals can spread between agents and outlive a context reset. The patch is a paragraph.

A 73-page preprint evolved instructions that jumped between coding agents and wrote themselves into the file that becomes the next system prompt. A short warning nearly stopped transmission.

Perspective Coverage

4 publishers
Builder
Builder 52%
Operator
Operator 39%
Investor
Investor 9%

Reality

Evidence68
Adoption
Insufficient
Hype gap+10
Incentives30
Confidence65
build4 publishers

Thomson Reuters priced the middle path at $40M, and still pays Anthropic

The $450,000 training run everyone is quoting is about one percent of the programme behind it. The in-house model took one CoCounsel feature; the agent layer under it stays licensed.

Perspective Coverage

4 publishers
Builder
Builder 35%
Operator
Operator 40%
Investor
Investor 25%

Reality

Evidence54
Adoption27
Hype gap+34
Incentives78
Confidence63