Skip to content

Topic

Storage-backed inference

Serving model weights from flash or SSD on demand, often through memory mapping, so a model larger than available RAM can run.

Current clusters