MoE models distribute parameters across k experts, reducing active experts' compute cost
Image: Unknown authorUnknown author, Public domain, via Wikimedia Commons
MoE models distribute parameters across k experts, reducing active experts' compute cost
Mixture of experts
Mixture of experts (MoE) divides problem space into homogeneous regions
the compute-optimal training ratio is: roughly 20 tokens per parameter
How can we train AI efficiently without wasting resources?
load balancing loss is needed in MoE
Can one expert handle all tasks perfectly?
AWQ does differently
AWQ selectively retains weights crucial for model performance, unlike traditional quantization
KV-cache reduces redundant computation in autoregressive generation
KV-cache stores previously computed outputs to avoid redundant calculations in autoregressive models
Tesla Model Y
Tesla Model Y is the world's best-selling electric vehicle in 2023
Swipe through 100 ML concepts daily
Open Pocket Polymath