Federal prosecutors charged EarthMade Computer owner Greg Lui with smuggling more than $300 million in GPU servers to China through Malaysia and Singapore. So far the case charges only Lui and treats the US manufacturers that sold him servers as parties he deceived with false end-user papers.
Perspective Coverage
20 publishers
- Builder
- Builder 21%
- Operator
- Operator 42%
- Investor
- Investor 37%
Reality
- Evidence68
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence65
CME plans to list two Nvidia GPU compute futures on October 5, pending regulatory approval, with prices running 36 months forward. Bitcoin miners that moved into AI hosting can hedge with them only if what they are paid moves with a GPU rental index.
Reality
- Evidence42
- Adoption10
- Hype gap+55
- Incentives78
- Confidence45
Thieves broke into two PlusAI trailers taken from Fremont, one Nvidia-branded, and found 40,000 pounds of test sand. Overhaul's tally of more than $150 million in AI cargo losses is small next to AI spending, but single thefts run into the millions, and that cost lands on whoever ships each load of chips.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+35
- Incentives65
- Confidence45
g factor's Qwen 3.8 27B benchmark has Together AI fastest at one stream, at 189.61 tok/s, while four of five engines finish within about 10% at 64 streams. Choosing a provider from these numbers starts with knowing how many streams the deployment will run at once.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+18
- Incentives72
- Confidence45
LessWrong post puts the GPU cost of an AI doing an hour of median human work at about 4 cents, against a $25 US median wage. Current API prices narrow that gap sharply, and for the hardest tasks they lift AI cost to the hourly rate of a skilled engineer.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence35
FastGPU's September 27 snapshot of 28 GPU clouds puts the cheapest hyperscaler H100 at $5.38 an hour, 3.0x the $1.79 market floor. That premium buys IAM, managed services and credits, and it is worth paying when a team actually uses them.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence50
Moonshot published 2.8 trillion open weights. At four bits per parameter that is about 1.4TB resident before any cache, which rules out the eight-way H100 node most teams assume.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence58
The orbital compute startup added an extension to its Series A largely to buy launch capacity before Falcon 9 retires in 2028, according to CEO Philip Johnston.
Perspective Coverage
5 publishers
- Builder
- Builder 34%
- Operator
- Operator 26%
- Investor
- Investor 40%
Reality
- Evidence55
- Adoption15
- Hype gap+45
- Incentives70
- Confidence60
NVIDIA's SWE-Serve scores the same 627 patches twice on 19 SGLang tasks, once with the live-serving tests and once without. The pass rate falls from 69.4% to 45.9%, and 242 of the 276 live tests came from SGLang itself.
Reality
- Evidence58
- Adoption30
- Hype gap−10
- Incentives70
- Confidence60
Ternary weights at 1.76 bits put a 27B model into a 5.9GB file and let a laptop decode it at 28.1 tokens a second. The retention figure comes from Prism ML's own benchmark suite, not the table on the model card.
Reality
- Evidence38
- Adoption42
- Hype gap+34
- Incentives74
- Confidence46
PyTorch says vLLM's frontier models now ship as hardware-specific flat definitions that torch.compile cannot trace, and the new HW agnostic layers are what users on other accelerators get instead. The overhead figure came from an H100.
Reality
- Evidence55
- Adoption45
- Hype gap+10
- Incentives60
- Confidence55
The Replicate build from hautechai takes one image URI and one query string per call, with both thresholds adjustable between 0 and 1. Throughput and memory are numbers you would have to measure yourself.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap−10
- Incentives40
- Confidence55
One 8,192-token session on a 27B model holds 512 MB of key-value cache, or 64 KB for every token generated. How many of those sessions fit in free VRAM sets serving concurrency, and paging decides the waste.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+24
- Incentives45
- Confidence48
kev packs a document and every typed question into one sequence and reads them in a single forward pass on a Mac. Its author scored four checkpoints against the real Jev on frozen items and published the reads that missed a pre-declared gate.
Publishers:scour.ing
Reality
- Evidence48
- Adoption12
- Hype gap−8
- Incentives55
- Confidence45
NVIDIA says multi-agent systems burn up to 15 times the tokens of a standard chat, and Nemotron 3 Super is its open-weight attempt to make each of those tokens cheaper to produce. The efficiency figures come with NVIDIA's own hardware and its own predecessor as the baselines.
Reality
- Evidence38
- Adoption18
- Hype gap+34
- Incentives88
- Confidence57
A dev.to walkthrough of a 600GB NVFP4 model on discounted 8xH100 spot nodes traces the two crashes that arrive before the first prompt to a pip resolver replacing numpy and a KV cache sized for a million tokens.
Reality
- Evidence26
- Adoption14
- Hype gap+34
- Incentives
- Insufficient
- Confidence28
A C4ADS investigation maps three routes for restricted accelerators into China, and almost all of the value it counted sits with a single importer whose ownership is still unresolved.
Reality
- Evidence52
- Adoption45
- Hype gap+28
- Incentives62
- Confidence50
An American Banker opinion piece argues that GPU-backed loans audit serial numbers and liens while the compute that repays them goes unmetered, and it points at power finance, where lenders advance against settled megawatt hours.
Reality
- Evidence36
- Adoption18
- Hype gap+33
- Incentives74
- Confidence38
An arXiv defense adds Gaussian noise to a served model's logits and solves for the scale that hits a chosen accuracy target, and it bounds how many repeated queries an attacker needs to average that noise away.
Reality
- Evidence44
- Adoption8
- Hype gap+12
- Incentives
- Insufficient
- Confidence42
AWS's NVRx walkthrough puts synchronous checkpointing ahead of GPU faults as the source of idle time on its FSDP jobs, and pairs a background save with a restart that re-enters training without cycling the container.
Reality
- Evidence54
- Adoption20
- Hype gap+18
- Incentives76
- Confidence56
Earlier coverage
- Block diffusion drafting collapses EAGLE-3's eleven serial forward passes into one
Build · September 15, 2026 · 1 publisher
- Robot makers buy the inference before they sell the robot
Invest · September 15, 2026 · 1 publisher
- Red Hat traces the inference bill to 140GB of memory reads per token
Product · September 14, 2026 · 1 publisher
- MLA's 512-Scalar KV Latent Cuts Cache 98%, Enabling 512 Concurrent 128k Streams
Build · September 13, 2026 · 1 publisher
- Nvidia weighs absorbing a quarter of GPU value loss to get lenders into AI factories
Invest · September 12, 2026 · 1 publisher
- Positron raises $875M for an inference chip built around commodity LPDDR5X instead of HBM
Science · September 10, 2026 · 1 publisher
- Huawei's 750,000 Ascend chips add up to under a twenty-fifth of Nvidia's 2026 compute
Invest · September 9, 2026 · 1 publisher
- OPAQUE ships a spec that keeps model weights sealed until the host proves what it is running
Security · September 9, 2026 · 1 publisher
- OPAQUE's Weight Custody Manifest lets the decryption key expire unless the runtime re-attests
Build · September 9, 2026 · 1 publisher
- A 6% driver reserve decides which models fit on a $2,000 pair of P40s
Build · September 8, 2026 · 1 publisher
- Nex-AGI's runnable N2.5 tiers ask for two H100s or sixteen H200s
Build · September 8, 2026 · 1 publisher
- Korea's AI buildout outspends its sovereign model program 2,600 to one
Invest · September 2, 2026 · 1 publisher
- Making protein folding agent-callable via Claude Science and BioNeMo NIM, with about 700 GB of storage needed
Build · August 31, 2026 · 1 publisher
- GPUThor turns GPU memory integrity into a tenancy question for anyone renting Ampere cards
Leadership · August 30, 2026 · 1 publisher
- a16z puts $1.1 billion behind racks drawing as much as fifty times legacy power
Invest · August 29, 2026 · 1 publisher
- An H100's MIG slices hand Chromium's WebGL straight back to the CPU rasteriser
Build · August 28, 2026 · 1 publisher
- a16z's $1.1bn hardware fund prices AI roadmaps in kilowatts per rack
Product · August 28, 2026 · 1 publisher
- Amap bounds long-horizon 3D mapping to a 12-frame window instead of a growing keyframe store
Build · August 28, 2026 · 1 publisher
- Databricks puts your training bill in the data loader and the checkpointer
Build · August 27, 2026 · 1 publisher
- Nvidia's August 26 print: 92% of the quarter rides on one segment
Invest · August 23, 2026 · 1 publisher
- Starcloud's $250M is a down payment on rockets it cannot yet buy
Build · August 21, 2026 · 2 publishers
- Anthropic's protein binders got tested by outside labs. The benchmark is still Anthropic's.
Product · August 19, 2026 · 1 publisher
- H100 rentals are back to $2.35 an hour, and your AI cost model is stale
Invest · August 16, 2026 · 1 publisher
- SpaceX's AI1 and the real test for orbital data centers: bottlenecks out minus bottlenecks in
Product · August 15, 2026 · 1 publisher
- A million satellites on paper: read the filings, not Musk's 2029 deadline
Invest · August 15, 2026 · 1 publisher