Skip to content

Topic

Model compression

Techniques such as quantisation, pruning and distillation that shrink trained neural networks so they need less memory and compute to run.

Current clusters