papersSEP 10 04:00 UTC
Paper argues FP8 with Ozaki Scheme II can substitute FP64 on next-gen NVIDIA GPUs
An updated arXiv preprint contends that low-precision FP8 matrix operations, when combined with the CRT-based Ozaki Scheme II error-compensation technique, can handle numerical workloads traditionally reserved for double-precision (FP64) hardware. The authors focus on NVIDIA's B300-class AI accelerators, claiming their tensor cores make this approach viable for a range of matrix-dominated scientific computing fields. The paper is the first part of a series challenging the assumption that dedicated FP64 units are essential for high-performance computing.