Live Feed/Together AI/Fact Record
Together AI logo
Together AI
feature 96% Confidence Gate September 10, 2026

To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

Together AI has ported the ThunderKittens library to support the NVIDIA Vera Rubin NVL72 architecture. The implementation includes a rebuilt NVFP4 GEMM kernel that achieves over 22 PFLOPS performance.

Verified State Diff

Comparison Mode:
- Previous State
ThunderKittens lacked native support and optimized NVFP4 GEMM kernels for the Vera Rubin NVL72 architecture.
+ Verified New State
ThunderKittens supports NVIDIA Vera Rubin NVL72 with a rebuilt NVFP4 GEMM kernel delivering over 22 PFLOPS.

Impact & Verification Analysis

WHO IS AFFECTED

AI infrastructure engineers, GPU kernel developers, and high-performance computing researchers.

WHY IT MATTERS

This demonstrates the ability to achieve near-native performance on next-generation NVIDIA hardware using custom kernels, reducing reliance on closed-source libraries like cuBLAS for high-throughput AI workloads.

Full Fact Overview

The integration involves a low-level optimization of the NVFP4 GEMM (General Matrix Multiply) specifically tailored for the Vera Rubin NVL72 hardware. By leveraging new ISA (Instruction Set Architecture) capabilities, the team improved efficiency from 42% of the theoretical roofline to a performance level competitive with NVIDIA's proprietary cuBLAS and CuTe DSL libraries.

Multi-Source Evidence Chain (1)

To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!Together AI
TRACKED ENTITY
Explore all historical Together AI changes
View Together AI Hub ➔