Skip to content

Topic

Inference serving frameworks

Software layers that batch, route and parallelise large-model requests across GPUs, including techniques such as disaggregated prefill and decode and expert parallelism.

Current clusters