April 2024
Intermediate to advanced
264 pages
6h 10m
English

Where does self-attention get its name, and how is it different from previously developed attention mechanisms?
Self-attention enables a neural network to refer to other portions of the input while focusing on a particular segment, essentially allowing each part the ability to “attend” to the whole input. The original attention mechanism developed for recurrent neural networks (RNNs) is applied between two different sequences: the encoder and the decoder embeddings. Since the attention mechanisms used in transformer-based large language models is designed to work on all elements of the same set, it is known as self-attention. ...
Read now
Unlock full access