v0.30.1rc0: [ROCm][CI] Add MI355 dense NVFP4 and MoRI kernel mirrors (#58281)
vLLM has introduced support for MI355 dense NVFP4 and MoRI kernel mirrors within the ROCm CI pipeline. This update enables specific hardware-accelerated precision formats for AMD Instinct MI355 accelerators.
Verified State Diff
Impact & Verification Analysis
Developers deploying LLMs on AMD Instinct MI355 hardware and infrastructure engineers managing ROCm-based inference clusters.
It accelerates the adoption of AMD hardware in high-performance inference environments by ensuring software-level compatibility with cutting-edge quantization formats and architectural optimizations.
Full Fact Overview
The integration of MI355 dense NVFP4 (NVIDIA Floating Point 4-bit) and MoRI (Mixture of Experts/ROCm-specific) kernel mirrors indicates a strategic alignment between vLLM and AMD's latest hardware roadmap. By adding these to the Continuous Integration (CI) pipeline, vLLM ensures that high-performance inference kernels for 4-bit precision and specialized MoE architectures are validated and optimized for the MI355 GPU architecture. This reflects a broader industry trend of adopting sub-8-bit quantization formats to maximize throughput on next-generation data center GPUs.