EXPERT

Tensor Cores, WMMA & Hardware Acceleration

Master Warp Matrix Multiply and Accumulate (WMMA), cuBLAS FFI bindings, and quantized FP8 matrix multiplication. Part of the free Rust Systems & AI Academy — every lesson below is open to everyone, no signup required.

2 lessons310 XP~38 min total100% free

// LESSONS IN THIS MODULE

  1. 01Warp Matrix Multiply and Accumulate (WMMA) Architecture20 min · 160 XP

    The Physics of NVIDIA Tensor Cores Standard CUDA cores execute scalar math (one FMA per clock). Tensor Cores execute entire matrix multiplications per...

  2. 02cuBLAS & Quantized FP8 Matrix Multiplication in Rust18 min · 150 XP

    Zero-Overhead cuBLAS Bindings While custom PTX is vital for unique activations, dense GEMM (General Matrix Multiply) operations are best delegated to...

Explore the full Rust Systems & AI Academy