AMD officially launched Helios, its first rack-scale AI system, at Advancing AI 2026 in San Francisco on July 23. The system combines 72 AMD Instinct MI455X GPUs with 6th-gen EPYC Venice CPUs and Pensando networking into a double-width rack (4OU) delivering 2.9 exaflops of FP4 inference performance. AMD claims Helios delivers 15% better compute performance and 50% greater high-bandwidth memory (HBM4) capacity than NVIDIA's Vera Rubin NVL72, while offering 30% more inference tokens per dollar.
Helios is priced between $5–5.5 million per rack and enters production immediately, with early customer deployments beginning in Q4 2026. Major customers announced include Microsoft Azure, Meta (1 GW committed to Helios), OpenAI (12 GW of AMD GPUs across OpenAI and Meta), Oracle, Anthropic (2 GW partnership with engineering collaboration to use Claude for ROCm optimization), and Tata Consultancy Services. Each rack provides 31 terabytes of HBM4 memory and 1.7 petabytes/second of aggregate bandwidth.
AMD's move directly challenges NVIDIA's >95% data center GPU market dominance; AMD currently holds ~4.5%. Helios represents AMD's most integrated play yet—bunching silicon, networking, and software with strategic partnerships on inference workloads (via Cerebras) and model optimization (via Anthropic, OpenAI). For architects planning multi-vendor AI infrastructure and evaluating total cost of ownership over time, AMD Helios signals credible rack-level competition in 2026–2027, with supply chains now diversifying beyond NVIDIA monopoly.