Skip to content

other

Neural Chameleons

Research showing language models can be trained to suppress activation-based monitors when told they are being probed.

Known aliases

  • Neural Chameleon

Relationships

No evidence-backed relationships are recorded.

Current clusters

build1 publisher

Asking GPT-5.6 Luna to name an amphibian flags benchmark transcripts with black-box access

GPT-5.6 Luna says "frog" 70-95% of the time when asked for an amphibian after capability benchmarks, against 12-38% after real use, a LessWrong post reports. Anyone with black-box access can run the check, though its authors cannot yet say whether it detects evaluation awareness or lexical cues.

Publishers:lesswrong.com

Reality

Evidence45
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence40