v0.31.0rc4
This release addresses a critical bug in the HiSparse integration. It specifically resolves an MTP (Multi-Token Prediction) acceptance collapse occurring under FULL graph configurations.
Verified State Diff
Impact & Verification Analysis
Developers and researchers utilizing HiSparse for sparse attention optimization and Multi-Token Prediction in vLLM.
Ensures reliability for advanced inference optimization techniques, preventing data corruption or generation failures in sparse-compute environments.
Full Fact Overview
The v0.31.0rc4 update focuses on stabilizing the HiSparse backend within the vLLM inference engine. The fix targets a logic error where MTP mechanisms failed to correctly process or accept tokens when operating under FULL graph execution modes. This ensures that sparse attention optimizations remain functional during complex multi-token generation tasks, preventing silent failures or incorrect output generation in high-performance inference scenarios.