Robust statistics

How can we estimate the center and spread of data even when it's messy?

Image: Shailaja.k, CC BY-SA 3.0, via Wikimedia Commons

Robust statistics

How can we estimate the center and spread of data even when it's messy?

Imagine you're trying to find the average height and variation in height of students in a school, but some students have unusual heights that don't fit the pattern.

We want to find a way to look at the heights that isn't thrown off by those few students who are much taller or shorter than everyone else. The technical term for this is "sufficient statistics."

Example

Let's say most students are 150 cm tall with a variation of 10 cm, but there's one student who's 190 cm tall. We still want to find the average height and variation that represents the typical student.

Remember this

Sufficient statistics help us estimate the true characteristics of a dataset without being misled by unusual values.

Related concepts

Swipe through 100 ML concepts daily

Open Pocket Polymath