Red Hat is proud to announce our results from the industry-standard MLPerf Inference v6. This submission builds on our track record across recent rounds: In v5. 1 , we demonstrated cost-effective Llama-3. 1-8B-FP8 inference with vLLM on NVIDIA H100 and L40S GPUs, and in v6. 0 we delivered results across Qwen3-VL, Whisper, and gpt-oss-120b on NVIDIA H200 and B200 and AMD Instinct MI350 GPUs, including the first Kubernetes-based submission.

1 results highlight 3 things: peak performance on Kubernetes, 1 inference engine (vLLM) spanning graphics processing units (GPUs) and central processing units (CPUs), and the viability of CPU-only serving for small models. Our submissions span 4 open-weight models—gpt-oss-120b, Qwen3-VL-235B-A22B, Llama-3. 1-8B, and Whisper-Large-v3—across NVIDIA GB200 NVL4 systems and CPU-only Intel® Xeon 6 servers.

We participate in MLPerf because it’s the industry's standardized, peer-reviewed measure of inference performance: It lets enterprises compare hardware and software stacks on equal footing, and it holds our open source stack accountable to the same bar as proprietary alternatives. These results illustrate Red Hat AI's ability to deliver competitive inference on Kubernetes infrastructure, across GPU and CPU architectures, using a consistent software layer.

Notably, 22 of 32 datacenter submitters in this round used vLLM somewhere in their stack, confirming its position as the industry-standard open source inference engine. The highest per-core throughput of all 2-socket CPU submissions The highest per-core throughput of all 2-socket CPU submissions Figure 1: Red Hat MLPerf Inference v6. 1 results, closed division, datacenter. GPT-OSS-120B is a 117-billion-parameter mixture-of-experts model for reasoning, agentic workflows, and code generation, with highly variable input lengths and strict server-scenario latency requirements.