Why can't we just guess probabilities correctly all the time?
Image: Unknown authorUnknown author, CC BY 4.0, via Wikimedia Commons
Why can't we just guess probabilities correctly all the time?
Imagine you're trying to guess the outcome of flipping a coin, but you're not sure if it's fair. Every time you guess, you want to know how good your guess was.
You want to measure how well your guesses match the actual outcomes. The better your guesses match, the smaller the difference between your guesses and reality.
Example
You guess heads 70 times out of 100 flips, but heads actually came up 50 times. Your guesses are off by 20%.
Remember this
The log-loss function helps you understand how far off your guesses are from the actual outcomes.
Text adapted from Wikipedia, licensed under CC BY-SA 4.0.
log-probabilities are used instead of probabilities: avoids numerical underflow
Why can't we just add up tiny chances over time?
cross-entropy equals negative log-likelihood for classification
Why does knowing the wrong probability help us measure information loss?
log-loss / cross-entropy loss penalizes: confident wrong predictions more heavily
How wrong predictions hurt more than they're supposed to
Entropy H = -Σ p(x) log₂ p(x) measures average surprise in bits
How do we measure uncertainty in everyday decisions?
Cross-entropy H(p,q) = -Σ p(x) log q(x) measures how well q approximates p
Ever wondered how well we can guess the outcome of a random event?
Fisher information
Fisher information measures information about unknown parameters
Swipe through more Machine Learning concepts
Open Pocket Polymath