
Why do Transformers sometimes seem to 'ignore' irrelevant parts of the input?
Image: Yuening Jia, CC BY-SA 3.0, via Wikimedia Commons
Why do Transformers sometimes seem to 'ignore' irrelevant parts of the input?
Imagine you're listening to a friend tell a story, but you're only interested in the part about the birthday party. You tune out the rest.
In a Transformer, there's a way to focus on parts of the input that matter for the task, like listening to the birthday party story, while ignoring the rest. This is called sparse attention.
Example
If there are 100 words in the story, the Transformer might only 'pay attention' to 10 words that are about the birthday party.
Remember this
Sparse attention reduces the amount of computation needed by focusing on fewer parts of the input.
Text adapted from Wikipedia, licensed under CC BY-SA 4.0.
Transformer (deep learning)
Transformers use multi-head attention for contextualizing tokens
ring attention does: distributes long sequences across multiple devices
How can a machine understand and generate human language?
most transformer operations are memory-bound, not compute-bound
Why do computers sometimes get tired?
Masking (behavior)
Can you not see what's right in front of you?
soft targets carry more information than hard labels: they encode class similarities
Why do some learning methods need to explore more than others?
Attention (machine learning)
Why do we sometimes zoom in faster on a scene?
Swipe through 100 ML concepts daily
Open Pocket Polymath