Red Hat releases NVFP4 Qwen3.8-Flash-Next checkpoint: MoE experts at 4-bit, rest BF16
Red Hat's AI division has uploaded a compressed checkpoint of Qwen3.8-Flash-Next to Hugging Face in which only the linear layers of the mixture-of-experts (MoE) modules are quantized to NVFP4, Nvidia's 4-bit floating-point format. The remainder of the model — router, embeddings, shared layers and all non-expert weights — stays in BF16. The architecture is Qwen4ExpForConditionalGeneration, accepting multimodal input (text, image, video) and producing text output. The checkpoint is tagged version 1.0 with an official release date of 27 August 2026 and is designed to run on vLLM with standard parameters: tensor-parallel-size 4, plus the Qwen3 tool-call-parser and reasoning-parser.
According to Red Hat's evaluations, the checkpoint achieves 99.1% accuracy recovery relative to the uncompressed model — defined as the compressed score divided by the original score, multiplied by 100, capped at 100%. For comparison, a parallel checkpoint from Inferact scores 97.6% and Nvidia's own scores 98.0%. Red Hat attributes the gap to differences in calibration data and observer implementation. All measurements were run through Inspect against a vLLM server using a single seed.
The model was produced with LLM Compressor using 1,024 samples drawn from the open-perfectblend dataset. Quantization was applied exclusively to the weights and activations of linear operators inside the MoE experts; no other components were altered. The result is a significant reduction in memory and disk footprint for the expert parameters — the heaviest part of an MoE model — without touching the rest of the graph. The launch command still requires four GPUs under tensor parallelism, indicating that VRAM demands remain non-trivial despite the compression.
The practical upside is memory and storage savings for the costliest segment of an MoE model while preserving near-original quality. Shipping a vLLM-ready checkpoint with explicit parameters lowers the barrier for teams that want to run a large multimodal model on a smaller hardware footprint. The 634 downloads recorded over the past month signal early interest rather than broad adoption.