Skip to content

Topic

Sparse mixture-of-experts models

Model architecture that routes each token through a small subset of many expert subnetworks, so total parameter count far exceeds the compute spent per token.

Current clusters