AMD Helios on Azure: Microsoft Relys on MI455X and Epyc Venice for AI Inference

Philipp Briel
Philipp Briel · 3 min. read

Microsoft Brings AMD Helios to the Cloud: The company will deploy AMD’s new Rackscale platform on Azure to run AI inference for its own services, Frontier models, and customers. Both companies announced this on July 20, 2026, as an expansion of their long-standing partnership—this time covering the full spectrum of GPUs, CPUs, networking, and software. Deliveries of Helios to customers, including Microsoft, will begin in the second half of 2026.

TL;DR

  • Microsoft will deploy AMD Helios on Azure for AI inference, launching in the second half of 2026.
  • Helios combines Instinct MI455X, Epyc “Venice” (Zen 6), Pensando networking, and ROCm.
  • It also includes two new Epyc VM series (HDv2, HXv2) plus the ND MI455X v7 GPU instance.
  • The deal is another blow in the race with Nvidia for AI data centers.

What’s Inside the AMD Helios Racks

Helios is AMD’s first integrated rack-scale solution—meaning no more individual accelerator cards that you plug into existing servers, but rather a complete system from a single source. A single rack houses up to 72 Instinct MI455X accelerators with a combined 31 TB of HBM4 memory and approximately 1.4 PB/s of bandwidth. AMD cites up to 2.9 FP4 exaFLOPS for inference and 1.4 FP8 exaFLOPS for training. This is complemented by Zen 6-based Epyc “Venice” CPUs, Pensando DPUs for networking, and the open software platform ROCm.

According to AMD, a single MI455X delivers 432 GB of HBM4 and up to 40 PFLOPS in FP4—the chip is clearly tailored for low-precision AI formats such as FP4, FP8, and BF16, i.e., typical inference operations. This is exactly where the market is shifting right now: away from pure training and toward application. As a result, CPUs are regaining importance.

Three New Azure Instances with AMD Technology

In addition to Helios, Microsoft is expanding its Azure portfolio with two CPU-only series based on Epyc Venice, as well as a GPU instance for AI inference:

Instance Intended Use Key Specifications
Azure HDv2 Agentic AI & Data Pipelines Nearly 500 CPU cores (Venice), 4 TB RAM, 32 TB NVMe, 400 Gbit network
Azure HXv2 Chip Design (EDA) 176 CPU cores (Venice, 3D V-Cache, >5 GHz), 2–4 TB RAM, 800 Gbit InfiniBand
ND MI455X v7 AI inference AMD Helios: Epyc Venice + Instinct MI455X + Pensando DPU

At the network layer, both companies are also further expanding their use of Pensando DPUs and integrating them with Azure Boost to scale throughput and efficiency across the entire Azure fleet.

Another move against Nvidia

“AMD and Microsoft have been building high-performance infrastructure together for years,” said AMD CEO Lisa Su. Microsoft CEO Satya Nadella emphasizes “performance, scalability, and freedom of choice” for customers. In other words: Azure does not want to rely solely on Nvidia. AMD is explicitly positioning Helios as a counterpart to Nvidia’s Vera Rubin platform—and the fact that a hyperscaler like Microsoft is adopting the solution “at scale” sends a strong signal. Lisa Su’s catchphrase, “Everyone wants Venice,” thus gains another prominent name.

Anyone who has been following AMD in the server space has been aware of this trend for some time—for example, from the expansion of the EPYC platforms among server manufacturers like TYAN. More details are expected to follow at AMD Advancing AI 2026, which begins on July 22. For Microsoft and AMD, however, the Azure deal is already one of their biggest joint steps to date in the AI data center space.

Sources: AMD Newsroom, Microsoft, AMD Instinct