Anthropic now watermarks Claude's text with a version of Google DeepMind's SynthID-Text to meet the EU AI Act's transparency rules. Only a holder of the secret key can measure the mark, a lean in token choice that says little until it spans many tokens.
Reality
- Evidence45
- Adoption40
- Hype gap0
- Incentives55
- Confidence45
Encoding replays an ordered merge list that was learned once, before training. Swap the list and the weights still load, but the model you tested is gone.
Reality
- Evidence74
- Adoption45
- Hype gap+5
- Incentives32
- Confidence66
The demo uploads a model file as base64 chunks over GET, then starts an inference server on it. It shows why egress and WAF rules keyed on the HTTP verb miss what the URL is doing.
Reality
- Evidence64
- Adoption9
- Hype gap+14
- Incentives30
- Confidence56
A sharptext.net column argues the loudest AI risk estimates should be scored against the record of the people issuing them. Of the four forecasts it assembles, only one can be checked inside a year.
Publishers:sharptext.net · vox.com Reality
- Evidence55
- Adoption30
- Hype gap+35
- Incentives65
- Confidence50
The offset in a J-lens readout mostly tracks how often a token appears, and scaling it by variance, after plain subtraction failed, lifted hidden-word elicitation to 0.805 from 0.665 on Gemma-2-9B-it. The paired test over 20 words gives p of about 0.19.
Reality
- Evidence46
- Adoption12
- Hype gap+10
- Incentives40
- Confidence56
A disputed million-dollar proof and a public resignation at Anthropic landed a day apart. Gebru's answer to WIRED puts the weight on how fast a lab's claim reaches a lawmaker with no way to check it.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence55
Shrijith Venkatramana's walkthrough of sampling puts the odds-ratio arithmetic behind temperature on the page. It shows how much of the difference between two runs of one prompt is settled after the model has finished computing.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+8
- Incentives30
- Confidence55
Hugo Vergnes reports 0.384 CORE from 65.3 billion tokens on eight rented B200s. The $998 buys 43 hours of node time, not the 139 GPU-hours of failed run that produced the recipe, and not the evenings that wrote the framework.
Reality
- Evidence52
- Adoption12
- Hype gap+22
- Incentives45
- Confidence62
Unitree's $66 billion valuation lost nearly half its value in a week, and robot brains still lack the data to justify the hype.
Reality
- Evidence44
- Adoption38
- Hype gap+42
- Incentives72
- Confidence48
Llama 3.1 was pretrained on 15 trillion tokens. A preteen manages on about 100 million words. Nobody knows how the child does it, and easily available web data may run out in the 2030s.
Reality
- Evidence54
- Adoption20
- Hype gap+12
- Incentives32
- Confidence52
Two models can quote identical per-token prices and still bill differently for the same string. The split is decided by merge tables you did not train and cannot assume.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+22
- Incentives30
- Confidence48