PewDiePie says OpenAI banned him twice while he used GPT-5.6 Sol outputs to train Ajax, a 9-billion-parameter model for home PCs. OpenAI has not commented publicly, but his account shows that a team training its own model on a frontier lab's answers can lose API access partway through the build.
Perspective Coverage
4 publishers
- Builder
- Builder 45%
- Operator
- Operator 35%
- Investor
- Investor 20%
Reality
- Evidence45
- Adoption5
- Hype gap+35
- Incentives60
- Confidence55
Cantina released Apex Flash-1, an open-weights security model fine-tuned on 50 of its own vulnerability cases. Its only performance evidence is a 60-task evaluation the company ran itself, and a separate build is modified to refuse fewer requests.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence40
Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
UW and Meta researchers report 59.4% on BrowseComp-Plus for a model that edits its own context, against 53.4% for Codex-style summarisation. The edited file is thrown away when the task ends, so memory that lasts across tasks is still the builder's job.
Reality
- Evidence42
- Adoption8
- Hype gap+12
- Incentives
- Insufficient
- Confidence45
Authors of a dev.to trainer comparison say summing task, cost and guardrail rewards can leave up to a third of GPU batches with zero gradients. They back per-channel normalization, as in GDPO, and want guardrails enforced as hard rules.
Reality
- Evidence40
- Adoption20
- Hype gap+30
- Incentives30
- Confidence40
AWS says running DeepEP over its EFA network on EKS gives mixture-of-experts reinforcement learning 40% more throughput. Whether that reaches another cluster depends on how much of each training step goes to expert traffic between nodes.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+30
- Incentives80
- Confidence35
One 8,192-token session on a 27B model holds 512 MB of key-value cache, or 64 KB for every token generated. How many of those sessions fit in free VRAM sets serving concurrency, and paging decides the waste.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+24
- Incentives45
- Confidence48
A dev.to walkthrough of DeepSeek's GRPO puts a 70B PPO training loop over 600GB of VRAM before any sharding. The critic-free version deletes the most expensive network and pays for it in sampled completions.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+30
- Incentives42
- Confidence45
An arXiv paper names the failure behavioral state decay and runs a side-car memory agent next to an unmodified action agent, reporting gains of 8.3 and 6.8 points of pass@1 on two long-horizon benchmarks.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+18
- Incentives60
- Confidence57
Retrieve-for-Train runs reinforcement learning once against a fixed corpus and distils the winning sub-query sets into a lightweight diffusion model. Adopting it puts a retrain schedule on whoever owns the catalogue.
Reality
- Evidence45
- Adoption8
- Hype gap+30
- Incentives60
- Confidence55
The band it joined is one where the best closed models still finish under 10% of tasks end to end at roughly $50 and 20 minutes each, so what post-training buys a law firm is price and hosting rather than capability.
Publishers:harvey.ai
Reality
- Evidence41
- Adoption
- Insufficient
- Hype gap+32
- Incentives76
- Confidence46
Greedy accuracy came back to baseline by 1,500 steps and per-sample correctness tripled, so the fall in pass@64 to 0.19 only surfaced when the authors paid for 64 samples a problem instead of one.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+18
- Incentives55
- Confidence45
Three LoRA stages on Qwen2.5-0.5B pushed reward to 1.0 and accuracy below the untrained base. The fix that recovered 43 points was an execution harness that runs both queries and compares the rows.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+18
- Incentives38
- Confidence55
Thinking Machines' own figures show that better prompts carried the models it tested from a coin flip to the mid-70s and then quit, leaving the last five points to an 80% trust bar to be bought with annotation time from the analysts being freed.
Publishers:thinkingmachines.ai
Reality
- Evidence34
- Adoption18
- Hype gap+38
- Incentives78
- Confidence46
A year of SFT and GRPO on 9B to 35B vision-language models, and the costliest bug never threw an exception. It supervised prose and scored a single letter.
Reality
- Evidence30
- Adoption12
- Hype gap−5
- Incentives32
- Confidence42
The model now writes its own training tasks and grading harnesses. That removes the bottleneck of hand-built tasks and replaces it with a harder one: rewards that cannot be gamed.
Reality
- Evidence38
- Adoption20
- Hype gap+24
- Incentives72
- Confidence54
An open-source stack pairs a deterministic Minecraft reimplementation with seed-level provenance, so a reinforcement learning result can be replayed instead of reconstructed.
Reality
- Evidence42
- Adoption10
- Hype gap+12
- Incentives34
- Confidence45
High-confidence "no problem" precision fell from a roughly 98% target to 90.8% one week past tuning. Closing alerts automatically buys you a monitoring job, not fewer analysts.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−10
- Incentives58
- Confidence55