Skip to content

Topic

Inference disaggregation

Serving architecture that runs the stages of a model request (media encoding, prefill, decode) as separate workers that batch, schedule and scale independently, exchanging intermediate state over an interconnect.

Current clusters