build1 publisher
NVFP4 and cache reuse cut MLPerf's edge agent workload to 24 minutes on one Jetson Thor
NVIDIA's submission runs the same Qwen3.6-27B as the llama.cpp reference on the same Jetson board and finishes 6.4x sooner. Most of the gap comes from prompt tokens the runtime never has to prefill.
Publishers:developer.nvidia.com
Reality
- Evidence58
- Adoption18
- Hype gap+20
- Incentives85
- Confidence62