Skip to content

other

Sparse mixture-of-experts

Neural network design in which a router activates only a small subset of many expert subnetworks for each token, keeping compute per token far below the model's total parameter count.

Known aliases

  • sparse MoE

Relationships

No evidence-backed relationships are recorded.

Current clusters