Skip to content

project

TensorRT-LLM

NVIDIA's open-source library for optimizing and serving large language model inference on NVIDIA GPUs, used in high-performance LLM deployment stacks.

Known aliases

  • NVIDIA TensorRT LLM
  • NVIDIA TensorRT-LLM
  • TensorRT LLM
  • TensorRT-LLM
  • TRTLLM
  • TRT-LLM

Relationships

No evidence-backed relationships are recorded.

Current stories

build1 publisher

Confidential inference on Blackwell retains 96-98 percent of throughput with CC-aware adaptations, NVIDIA reports

NVIDIA measured TensorRT LLM holding 96.1 to 98.2 percent of its non-confidential output throughput on Blackwell, and it got there by unpinning host memory on the affected paths, moving decode readback off the scheduler thread, and timing kernel tactics with the GPU's global timer instead of CUDA events.

Reality

Evidence58
Adoption22
Hype gap+10
Incentives80
Confidence55
build5 publishers

Alibaba ships the Qwen4 architecture as open weights before the flagship exists

Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.

Perspective Coverage

5 publishers
Builder
Builder 52%
Operator
Operator 28%
Investor
Investor 20%

Reality

Evidence58
Adoption52
Hype gap+32
Incentives76
Confidence71