Skip to content

Topic

Inference memory footprint

The accelerator memory a served model consumes for its weights, cache and activations, and the formats used to reduce it.

Current clusters