2022 in science

Why do Transformers sometimes seem to 'ignore' irrelevant parts of the input?

Image: Yuening Jia, CC BY-SA 3.0, via Wikimedia Commons

2022 in science

Why do Transformers sometimes seem to 'ignore' irrelevant parts of the input?

Imagine you're listening to a friend tell a story, but you're only interested in the part about the birthday party. You tune out the rest.

In a Transformer, there's a way to focus on parts of the input that matter for the task, like listening to the birthday party story, while ignoring the rest. This is called sparse attention.

Example

If there are 100 words in the story, the Transformer might only 'pay attention' to 10 words that are about the birthday party.

Remember this

Sparse attention reduces the amount of computation needed by focusing on fewer parts of the input.

Related concepts

Swipe through 100 ML concepts daily

Open Pocket Polymath