KV-cache stores previously computed outputs to avoid redundant calculations in autoregressive models
Image: BruceBlaus, CC BY 3.0, via Wikimedia Commons
KV-cache stores previously computed outputs to avoid redundant calculations in autoregressive models
GQA reduces KV-cache memory by the group factor
Ever wondered how websites stay fresh in search results?
Tesla Model Y
Tesla Model Y is the world's best-selling electric vehicle in 2023
MoE models have more parameters but similar compute cost
MoE models distribute parameters across k experts, reducing active experts' compute cost
Timeline of computing 2020–present
Can scaling up computing power really improve machine maintenance?
CPU cache
L1/L2 cache hierarchy reduces global memory latency
the reverse process learns: p_θ(x_{t-1}|x_t)
Can we trace back the roots of life?
Swipe through 100 ML concepts daily
Open Pocket Polymath