Satlyt raised an $8 million seed round for software that runs AI models on satellites built by other companies. For operators, the part that has flown so far is onboard triage that cuts downlink costs, and the multi-satellite cloud in the founder's pitch has yet to be tried.
Reality
- Evidence50
- Adoption22
- Hype gap+35
- Incentives70
- Confidence55
Featherless open-sourced Simple Jev, a library that has open models pick fixed labels or yes/no answers, with hosting from $0.03 per million input tokens. CEO Eugene Cheah pitches it for classification work teams now pay frontier-model prices to run.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+35
- Incentives75
- Confidence45
Misspelling 70% of a prompt's words left Claude models' scores unchanged across about 4,900 test sessions, but one wrong punctuation mark cost 8 to 23 points. Both breaks that stuck erased the line between instruction and data, so delimiters deserve the review time that spelling gets.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
Jędrzej Maczan's paper finds the chat template turns on the 'just an AI' disclaimer in all eight open instruct models he tested. For eval teams, a model's self-description now depends on a formatting step most developers never see.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
A Kaggle-challenge benchmark called ART scores models on whether they still flag a function after the fix is applied. On eight synthetic pairs, the difference between price tiers showed up only on the patched half.
Reality
- Evidence47
- Adoption12
- Hype gap−5
- Incentives58
- Confidence44
Every local runner now reads the same GGUF file, so the binding decision is the quant tag and the gigabyte or two of context that has to fit beside it. Ollama sets the GPU offload itself; llama.cpp lets you set it.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+24
- Incentives44
- Confidence45
Serge Kernbach's dev.to build notes put a 9B to 27B assistant on two RTX 4070s totalling 24 GB and argue that integration beats raw model quality. The figures he publishes are power draw and PCIe bandwidth.
Reality
- Evidence32
- Adoption12
- Hype gap+30
- Incentives22
- Confidence42
A University of Washington study of public WildChat prompts finds fiction in more than a third of ChatGPT interactions and in 7 percent of users, against the 1.4 percent OpenAI's own research reports.
Reality
- Evidence55
- Adoption62
- Hype gap+15
- Incentives45
- Confidence55
A proposed scaling law for distillation says the winner flips on two things: whether you already own a teacher, and how many students you intend to serve.
Reality
- Evidence68
- Adoption34
- Hype gap+24
- Incentives44
- Confidence64
The headline confirm rate scored agreement with the scanner's claim and bug detection in one number, so the follow-up reruns the same 200 OWASP slices with the flag removed and the predictions committed first.
Reality
- Evidence48
- Adoption15
- Hype gap+8
- Incentives35
- Confidence45
A researcher ran the same FreeBSD scan through base and abliterated open-weight builds and found the uncensored ones graduating three to four times as many findings, while the most aggressive one never surfaced the actual CVE.
Publishers:clearbluejar.github.io
Reality
- Evidence58
- Adoption14
- Hype gap+14
- Incentives22
- Confidence46
Nvidia says the Hub will stay accelerator-agnostic and that its own compute will never be required, a claim that cannot be checked against behaviour until well after a close expected in the first half of 2027.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+22
- Incentives74
- Confidence56
The catalog serves its single-page app shell for malformed requests and for throttling alike, so the status line tells you nothing. The only gate that holds under load is checking the content type before you parse.
Reality
- Evidence52
- Adoption10
- Hype gap−12
- Incentives25
- Confidence58
Half of those rejections say the same thing, that the proposing model pushed severity past the CVSS evidence it had just cited. That is a real finding about the proposer, and it still says nothing about whether a same-family reviewer would have caught it.
Reality
- Evidence44
- Adoption8
- Hype gap+12
- Incentives48
- Confidence38
The block from U+E0000 to U+E007F renders as nothing and pastes as whitespace, so the only reader is the model you sent to the page. The control that held was refusing to let that model pick a recipient.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+22
- Incentives72
- Confidence58
An arXiv tracing study of Claude Code agents on Gemma and Qwen measured prefix-cache hit rates between 84.6 and 99.5 percent, which moves the serving bottleneck to how long you can keep KV blocks resident between tool calls.
Reality
- Evidence58
- Adoption25
- Hype gap+12
- Incentives
- Insufficient
- Confidence55
Google's Gemma milestone arrives with an engineer's caveat and no breakdown. Alibaba's rival claim of 3bn Qwen downloads is about 1.5 times what Hugging Face independently counted.
Reality
- Evidence34
- Adoption61
- Hype gap+38
- Incentives79
- Confidence52
llama.cpp will not quantize a V cache without Flash Attention. Which half of the KV cache you can still shrink, and whether you measured it or guessed it, sets the context you can actually ship.
Reality
- Evidence58
- Adoption24
- Hype gap+12
- Incentives34
- Confidence46
The acc and acc_norm split in lm-eval-harness can move in opposite directions on one checkpoint. Pick the metric before you train, and say which one you picked.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+14
- Incentives18
- Confidence46
China's AI grouping went from 29 signatories to 38 while the State Department prepared a warning against "duplicative initiatives". Stack choices are becoming jurisdictional.
Reality
- Evidence44
- Adoption61
- Hype gap+17
- Incentives63
- Confidence48