
How can we estimate the center and spread of data even when it's messy?
Image: Shailaja.k, CC BY-SA 3.0, via Wikimedia Commons
How can we estimate the center and spread of data even when it's messy?
Imagine you're trying to find the average height and variation in height of students in a school, but some students have unusual heights that don't fit the pattern.
We want to find a way to look at the heights that isn't thrown off by those few students who are much taller or shorter than everyone else. The technical term for this is "sufficient statistics."
Example
Let's say most students are 150 cm tall with a variation of 10 cm, but there's one student who's 190 cm tall. We still want to find the average height and variation that represents the typical student.
Remember this
Sufficient statistics help us estimate the true characteristics of a dataset without being misled by unusual values.
Text adapted from Wikipedia, licensed under CC BY-SA 4.0.
to standardize: when you need zero mean and unit variance for gradient-based optimization
Why do we need to make data uniform before training a model?
the multivariate Gaussian is parameterized by: mean vector μ and covariance matrix Σ
Ever wondered how weather patterns or stock market trends can be predicted with surprising accuracy?
Bootstrapping (statistics)
Small sample sizes can mislead standard error estimates
the Dirichlet distribution does: distribution over probability simplices
How do we predict the likelihood of various outcomes in uncertain situations?
Resampling (statistics)
Bootstrapping samples with replacement to estimate distributions
Markov chain Monte Carlo
MCMC samples from complex posterior distributions
Swipe through 100 ML concepts daily
Open Pocket Polymath