Adam adjusts learning rates per-parameter, SGD generalizes better with tuning
Image: Official GDC, CC BY 2.0, via Wikimedia Commons
Adam adjusts learning rates per-parameter, SGD generalizes better with tuning
Adam combines momentum and RMSprop: adapts per-parameter learning rates
Ever wondered how a car adjusts its speed for smooth turns?
Neural network (machine learning)
Ever tried adjusting the learning rate like tuning a musical instrument?
the β₁ and β₂ hyperparameters control in Adam
β₁ controls the exponential decay rate of the first moment estimates; β₂ controls the exponential decay rate of the second moment estimates in Adam optimizer
batch size affects generalization: larger batches find sharper minima
Larger batch sizes lead to sharper minima, enhancing generalization by providing more accurate gradient estimates
MoE models have more parameters but similar compute cost
MoE models distribute parameters across k experts, reducing active experts' compute cost
data augmentation does for generalization: artificially expands training set
How can you teach a computer to see better?
Swipe through 100 ML concepts daily
Open Pocket Polymath