Skip to content

Topic

GPU Memory Budgeting

Sizing model weights, context windows and KV caches against the VRAM a device actually makes available after driver and runtime overhead, rather than the capacity printed on the card.

Current clusters