Skip to content

Topic

Model weight caching

Keeping model checkpoints on storage attached to the compute node so serving processes read them locally instead of fetching them over the network.

Current clusters