Skip to content

project

Mooncake

KVCache-centric disaggregated LLM serving architecture, published with request traces that are reused as a benchmarking workload format.

Current clusters