Skip to content

other

Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

2017 Google Brain paper that introduced sparse top-k gating for mixture-of-experts layers, along with the auxiliary loss used to keep expert usage balanced.

Known aliases

  • Shazeer et al. 2017
  • Sparsely-Gated Mixture-of-Experts Layer

Relationships

No evidence-backed relationships are recorded.

Current clusters