Tensor cores perform 4x4 matrix multiply using optimized GEMM (General Matrix Multiply) instructions in one clock cycle
Image: FMNLab, CC BY 4.0, via Wikimedia Commons
Tensor cores perform 4x4 matrix multiply using optimized GEMM (General Matrix Multiply) instructions in one clock cycle
tensor cores are
Why can computers crunch numbers faster than humans?
tl.dot does in Triton: block-level matrix multiply using tensor cores
tl.dot performs block-level matrix multiplication using tensor cores in Triton
Nvidia
Ever wondered how video games run so smoothly on your computer?
quantization to INT8 doubles throughput
Quantization to INT8 doubles throughput because tensor cores process INT8 2x faster
Matrix multiplication algorithm
Ever wondered how computers speed up multiplying huge numbers?
instruction-level parallelism (ILP) achieves: multiple operations per clock cycle
Ever wondered how computers can do so many tasks at once?
Swipe through 100 ML concepts daily
Open Pocket Polymath