Skip to content

Topic

Time to first token

Latency measure for generative models covering the delay between a request arriving and the first output token being returned.

Current clusters