
TensorRT optimizes deep learning inference by quantizing and fusing operations for NVIDIA GPUs
Image: BigRiz, CC BY-SA 3.0, via Wikimedia Commons
TensorRT optimizes deep learning inference by quantizing and fusing operations for NVIDIA GPUs
Nvidia
Ever wondered how video games run so smoothly on your computer?
tensor cores are
Why can computers crunch numbers faster than humans?
quantization to INT8 doubles throughput
Quantization to INT8 doubles throughput because tensor cores process INT8 2x faster
the compute-optimal training ratio is: roughly 20 tokens per parameter
How can we train AI efficiently without wasting resources?
XLA does for TensorFlow/JAX: compiles computation graphs for TPU/GPU execution
How can you speed up your favorite video game?
Tensor network
Ever wondered how scientists manage massive data without endless storage?
Swipe through 100 ML concepts daily
Open Pocket Polymath