Skip to content

Topic

AI inference cold start

The delay between launching a model-serving replica and having it ready to answer requests, dominated for large models by loading weights into GPU memory.

Current clusters