Transformers
Topic view
Transformers
Current observations and durable explanations connected to Transformers.
Current observations
Discovery
Transformers
Looped Transformers versus traditional depth
Transformers
Autoregressive models, step by step
Durable context
Notes
Transformers
Attention mechanism in detail
A step-by-step account of queries, keys, values, masking, scaling, softmax, and the final weighted sum.Transformers
Multi-head attention
How parallel attention projections create complementary views and recombine them into one representation.Transformers
Positional encoding and the encoder-decoder
Where position, masks, residual paths, encoder blocks, and decoder blocks fit in a complete Transformer.