Can we truly measure how good a machine translation is?
Image: Sora / OpenAI, Public domain, via Wikimedia Commons
Can we truly measure how good a machine translation is?
Imagine you're trying to understand a foreign movie without subtitles. You rely on a translation app, but it doesn't always get the meaning right.
The BLEU score helps us figure out how close the app's translations are to what a human would say. It's like comparing two different paths to the same destination.
Example
If a human translator says "The quick brown fox jumps over the lazy dog," and the app says "The fast brown fox leaps over the sleepy dog," we can use BLEU to see how similar they are.
Remember this
The BLEU score tells us how closely the machine-generated translation matches a human's translation.
Text adapted from Wikipedia, licensed under CC BY-SA 4.0.
BLEU
Ever wondered how computers know if a translation makes sense?
word error rate (WER) measures: edit distance between predicted and reference transcriptions
Ever wondered how machines understand speech as we do?
weight tying does in language models: shares embedding and output projection matrices
Ever wonder how machines understand the sequence of words in a sentence?
BLEU vs ROUGE: BLEU measures precision of n-grams, ROUGE measures recall
BLEU measures precision of n-grams, ROUGE measures recall
Large language model
LLMs can generate, summarize, translate, and analyze text in many contexts
ring attention does: distributes long sequences across multiple devices
How can a machine understand and generate human language?
Swipe through 100 ML concepts daily
Open Pocket Polymath