Paged attention (vLLM) improves serving throughput by reducing latency through non-contiguous KV-cache pages, enabling faster data retrieval
Image: Daniel Voigt Godoy, CC BY 4.0, via Wikimedia Commons
Paged attention (vLLM) improves serving throughput by reducing latency through non-contiguous KV-cache pages, enabling faster data retrieval
GQA reduces KV-cache memory by the group factor
Ever wondered how websites stay fresh in search results?
grouped query attention (GQA) does
GQA shares KV heads across multiple Q heads for efficient parameter usage
LSM trees optimize: write-heavy workloads by buffering writes in memory
Ever wondered how your favorite social media app handles millions of new posts every minute?
Flashbulb memory
Flashbulb memories are vivid but not always accurate
CPU cache
L1/L2 cache hierarchy reduces global memory latency
Attention (machine learning)
Why do we sometimes zoom in faster on a scene?
Swipe through 100 ML concepts daily
Open Pocket Polymath