Skip to content

Topic

Transformer attention

The operation by which a transformer weights every input token against the others, and the origin of most limits on how much context a model can use well.

Current clusters