EXPERT
Tensor Cores, WMMA & Hardware Acceleration
Master Warp Matrix Multiply and Accumulate (WMMA), cuBLAS FFI bindings, and quantized FP8 matrix multiplication. Part of the free Rust Systems & AI Academy — every lesson below is open to everyone, no signup required.
2 lessons310 XP~38 min total100% free
// LESSONS IN THIS MODULE
- 01Warp Matrix Multiply and Accumulate (WMMA) Architecture20 min · 160 XP
The Physics of NVIDIA Tensor Cores Standard CUDA cores execute scalar math (one FMA per clock). Tensor Cores execute entire matrix multiplications per...
- 02cuBLAS & Quantized FP8 Matrix Multiplication in Rust18 min · 150 XP
Zero-Overhead cuBLAS Bindings While custom PTX is vital for unique activations, dense GEMM (General Matrix Multiply) operations are best delegated to...