Skip to content

Topic

Inference throughput

How fast a served model emits tokens, and the batch size, output length, hardware configuration and measurement method that determine any published tokens-per-second figure.

Current clusters