Skip to content

Topic

Attention Architectures

Design variants of transformer attention and the trade-offs each makes between cached state per token and representational capacity.

Current clusters