build1 publisher
Confidential inference on Blackwell retains 96-98 percent of throughput with CC-aware adaptations, NVIDIA reports
NVIDIA measured TensorRT LLM holding 96.1 to 98.2 percent of its non-confidential output throughput on Blackwell, and it got there by unpinning host memory on the affected paths, moving decode readback off the scheduler thread, and timing kernel tactics with the GPU's global timer instead of CUDA events.
Publishers:developer.nvidia.com
Reality
- Evidence58
- Adoption22
- Hype gap+10
- Incentives80
- Confidence55