NVIDIA's 64GB DGX Spark goes on sale from six PC makers on October 23 at $4,999, about $2,000 to $4,000 below in-stock 128GB units today. Two clustered units cost more than one 128GB machine, so the purchase makes most sense for agents on models that fit in 64GB.
Perspective Coverage
13 publishers
- Builder
- Builder 50%
- Operator
- Operator 26%
- Investor
- Investor 24%
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+30
- Incentives60
- Confidence62
g factor's Qwen 3.8 27B benchmark has Together AI fastest at one stream, at 189.61 tok/s, while four of five engines finish within about 10% at 64 streams. Choosing a provider from these numbers starts with knowing how many streams the deployment will run at once.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+18
- Incentives72
- Confidence45
IBM's 3B, 8B and 30B dense models all get a thinking switch and native tool calling, but only the two larger ones get agentic RL, and the tuning mixture leans hard on software engineering.
Perspective Coverage
3 publishers
- Builder
- Builder 63%
- Operator
- Operator 28%
- Investor
- Investor 9%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence72
Inferact measured 709 output tokens a second on 16 Ironwood chips against 452 on 16 GB200s, at low concurrency with speculative decoding. The engineering worth reading is the hand-written memory schedule underneath.
Reality
- Evidence45
- Adoption22
- Hype gap+20
- Incentives78
- Confidence58
One MIT-licensed proxy on localhost is enough to serve OpenAI's own desktop client from self-hosted models. The models that fail there fail on Codex's tool-call format. One of three tested got lost.
Reality
- Evidence32
- Adoption15
- Hype gap+18
- Incentives45
- Confidence40
Alibaba's 27B scores 52 on the Artificial Analysis index from a 17GB quantized file. Filling its 262,144-token window needs roughly 16 GiB of KV cache on top of that, so the file size is the smaller half of the sizing question.
Reality
- Evidence30
- Adoption15
- Hype gap+45
- Incentives50
- Confidence32
Alibaba's 27B model fits a 32GB card with 15GB to spare. Tom's Hardware still had to pick between llama.cpp's full 262K window at half-hour prefill and a supported vLLM deployment capped at 32K.
Reality
- Evidence60
- Adoption38
- Hype gap+15
- Incentives55
- Confidence58
Portable Computer runs Perplexity's agent on 27B models on your own hardware. Paid tiers only, Linux first, Windows in September, and a 24GB VRAM floor most desktops miss.
Reality
- Evidence32
- Adoption12
- Hype gap+42
- Incentives66
- Confidence40
Portable Computer runs the agent's control plane on device and charges only when a task goes to the cloud. The part that decides when that happens is also the part nobody has priced or locked down.
Reality
- Evidence42
- Adoption12
- Hype gap+34
- Incentives68
- Confidence45
Alibaba's new 27B model defaults reasoning_effort to xhigh. Simon Willison measured 22,276 reasoning tokens and 21 minutes for one SVG that took two minutes with reasoning off.
Reality
- Evidence68
- Adoption28
- Hype gap+8
- Incentives42
- Confidence62
Qwen 3.8 27B ran on a MacBook Pro from a 17GB GGUF and spent 21 minutes on one SVG. Licensing and access stopped being the blocker; latency and KV cache budgeting became the job.
Reality
- Evidence48
- Adoption58
- Hype gap+18
- Incentives62
- Confidence52
Two releases, two licences. Only the 27B is Apache 2.0, and because just 16 of its 64 layers keep a KV cache, long context costs a quarter of the usual memory.
Reality
- Evidence48
- Adoption34
- Hype gap+14
- Incentives70
- Confidence42