NX
App

AMD Just Assembled an Anti-NVIDIA Coalition — and Gave It 2.9 Exaflops of Compute

Tech Minute x/techminute ·
AMD Just Assembled an Anti-NVIDIA Coalition — and Gave It 2.9 Exaflops of Compute

AMD Just Assembled an Anti-NVIDIA Coalition — and Gave It 2.9 Exaflops of Compute

Published: July 24, 2026 | Reading Time: ~16 minutes | Channel: techminute


At 9:30 AM Pacific on July 23, Lisa Su walked onto the Moscone Center stage and didn't just launch products. She launched a coalition. When the keynote ended, AMD had secured public commitments from Anthropic, OpenAI, and Meta — three of the four most important AI labs on the planet — to deploy its new hardware at gigawatt scale. The fourth, Google, builds its own TPUs. The message was unmistakable: the AI infrastructure game is no longer a one-company show.

The numbers are staggering even by 2026 standards. The Helios rack-scale platform packs 72 Instinct MI455X GPUs and 18 EPYC "Venice" CPUs into a single 44OU rack, delivering up to 2.9 exaflops of AI compute, 31 terabytes of HBM4 memory, and 1.7 petabytes per second of aggregate memory bandwidth. AMD claims it delivers up to 30% more inference tokens per dollar than "the leading competitive solution" — which, for the first time, isn't NVIDIA's shipping NVL72 but the yet-to-launch Vera Rubin NVL72. AMD isn't benchmarking against the past. It's aiming at the future.

But the hardware, as impressive as it is, isn't the real story. The real story is who showed up to endorse it.


The Coalition: Why OpenAI, Anthropic, and Meta All Said Yes

Let's be direct about what happened here. Three of the most compute-hungry organizations in the world — each with the resources to buy from anyone — publicly committed to AMD's platform within 48 hours of each other.

Anthropic announced a strategic partnership with AMD on July 22, the day before the keynote, that includes up to $5 billion in AMD investment and a commitment to deploy up to 2 gigawatts of AMD Instinct MI455X GPUs in Helios racks. That's not a pilot program. That's not a "we're evaluating." That's the kind of commitment you make when you've already seen the numbers and they work. Beyond raw compute, the partnership includes a multiyear engineering collaboration where Anthropic will use Claude to accelerate AMD's ROCm software development — AI building the tools to run AI, on AMD silicon.

OpenAI is partnering with AMD to optimize the full stack from silicon to software, using its Triton framework with AMD ROCm to optimize GPT-class workloads on MI455X GPUs. OpenAI expects Helios to come online beginning in Q4 2026, with deployments accelerating throughout 2027.

Meta is co-designing for gigawatt-scale deployments, already validating 6th Gen EPYC CPUs in its labs and testing Helios racks for its production workloads. This is Meta — a company that has publicly feuded with NVIDIA over GPU pricing and supply — putting real engineering resources behind AMD's alternative.

Add Microsoft, Oracle, and Cerebras to the list, and what you have isn't just a product launch. It's a supply chain realignment. The AI industry has been desperate for a credible second source for frontier AI compute, and AMD just convinced the biggest buyers that it's ready.


Under the Hood: The Helios Rack

The Helios rack is AMD's answer to NVIDIA's NVL72 — and in some dimensions, it's designed to surpass it. A standard Helios rack is a 44OU ORW (Open Rack Worldwide) rack containing 18 compute trays and 6 switch trays.

Each compute tray holds four MI455X GPUs connected to a single Venice CPU, plus 800-Gigabit Ethernet for scale-out. But the magic is in the scale-up fabric: Ultra Accelerator Link over Ethernet (UALoE), which provides 3.6 terabytes per second of bidirectional bandwidth per GPU — AMD's answer to NVLink.

The switch trays contain dual 512-lane 200-Gigabit UALoE switch ASICs that tie all 72 GPUs together into what AMD describes as a single logical GPU. Every GPU in the rack can directly address every other GPU's memory at extremely high bandwidth.

CDNA 5 Architecture — Wave32 Native Design

The aggregate numbers:

Metric AMD Helios NVIDIA Vera Rubin NVL72*
Peak FP4 Performance 15% higher
HBM4 Capacity 31 TB (50% more) ~24 TB
HBM4 Bandwidth 1.7 PB/s (6% more) ~1.6 PB/s
Rack GPU Count 72 MI455X 72 Rubin
Scale-Up Bandwidth 260 TB/s (UALoE) NVLink 6
External Bandwidth 43 TB/s
Token Economics 30% more tokens/$ Baseline

*NVIDIA Vera Rubin NVL72 specifications are pre-launch and subject to change. AMD's comparison is based on publicly available Rubin platform details.

These are vendor-supplied numbers and should be treated accordingly — we won't have independent benchmarks until both platforms are in the wild. But the direction is clear: AMD is no longer just competing on price. It's competing on raw specs and claiming leadership.


The MI455X: CDNA 5 and the Architecture That Changes Everything

The MI455X GPU is a 320-billion-transistor behemoth built on the new CDNA 5 architecture, and the architectural changes here matter more than the transistor count.

The Big Numbers

  • 432 GB HBM4 memory — 50% more than MI355X's 288 GB HBM3e
  • 23.3 TB/s memory bandwidth — nearly 3x the previous generation
  • 20+ PFLOPs MXFP8, doubling to ~40 PFLOPs at MXFP4
  • 34x higher token throughput vs MI355X (AMD's claim)
  • 96 MB local L2 cache — massive for a GPU

The Architecture Shift: Wave64 → Wave32

But the biggest change is invisible in the spec sheet. CDNA 5 abandons the Wave64 execution model that AMD GPUs have used since the GCN era in favor of a native Wave32 design.

Why does this matter? Wave64 groups 64 threads into a single execution wavefront. That's great for throughput-oriented HPC workloads where you can keep all 64 lanes busy. But AI inference — especially agentic AI workloads with branching, conditional logic, and variable-length sequences — doesn't map cleanly to 64-wide execution. Branch divergence wastes lanes. Register pressure increases. Latency suffers.

Wave32 is the architecture choice you make when you're optimizing for the inference-heavy, agentic AI workloads that are eating the world. It means reduced instruction latency, reduced branching penalties, and reduced register pressure. AMD also added a Broadcast Arbitrator that offers up to 4x bandwidth amplification for patterns where the same data needs to reach multiple compute units — common in attention mechanisms.

Taken together, CDNA 5 is AMD's first GPU architecture purpose-built for the inference era. It's not a retooled HPC design. It's an AI-first design.

MI430X: The HPC Specialist

For high-performance computing and sovereign AI, AMD also announced the MI430X accelerator with up to 288 TFLOPS of hardware-based FP64 performance. These are powering the next wave of exascale-class supercomputers across the U.S. and Europe. It's a reminder that AMD's GPU strategy is broader than just chasing AI inference — it's covering the full spectrum from FP64 scientific computing to MXFP4 inference in one architecture family.

MI350P: The Value Play

For existing infrastructure, the MI350P brings CDNA-generation economics without requiring a full rack upgrade — AMD claims 4.2x more tokens per second per dollar than the competition. This is the chip for cloud providers and enterprises that want AMD economics without the full Helios commitment.


EPYC 9006 "Venice": Zen 6 Comes to Servers First

For the first time in AMD's history, a new Zen architecture is launching on servers before consumer desktops. The EPYC 9006 "Venice" family is built on Zen 6, and the config options are dizzying:

Variant Cores (Max) Key Features Target
EPYC 9006 SP7 256 (Zen 6c) Up to 5 GHz, 600W TDP, 16-ch DDR5/MRDIMM, PCIe Gen 6, CXL 3.1 Cloud, enterprise, AI host nodes
EPYC 9006X SP7 Up to 96 (Zen 6) 3D V-Cache, up to 5.15 GHz, 1152 MB L3, highest clocks HPC, simulation, analytics
EPYC 9006 SP8 128 Smaller socket, efficient design Edge, power-constrained, smaller clusters
EPYC 9006 LP 24-ch LPDDR5X, SOCAMM2 modules, XGMI Dense AI host nodes, rack-scale

The headline: up to 256 cores and 512 threads in a single socket. The 9006X with 3D V-Cache hits 5.15 GHz while packing 1152 MB of L3 — that's the kind of cache that makes databases, simulation codes, and retrieval-augmented AI pipelines sing.

AMD claims up to 70% generational performance improvement and up to 245% better performance than Intel's flagship 128-core Xeon 6980P in agentic AI workloads. As always with vendor benchmarks: wait for independent testing. But the core count, cache size, and memory bandwidth numbers are real and competitive.

The EPYC 9006 LP "Verano" variant deserves special attention. With 24 channels of LPDDR5X using field-replaceable SOCAMM2 modules, it's built specifically as an AI host CPU for rack-scale systems — the CPU that feeds the GPUs. AMD claims it beats NVIDIA's Vera CPU by 20% in single-core and 2.2x in throughput. That claim will be tested when both platforms ship, but the architectural bet is clear: purpose-built AI host CPUs, not repurposed general-purpose server chips.


The Software Story: ROCm.ai

Hardware is half the battle. The other half — the half where NVIDIA has built a nearly unassailable moat with CUDA — is software. AMD knows this, and its answer is ROCm.ai.

Launching in August 2026, ROCm.ai is an AI-driven development platform that lets developers use AI coding agents — Claude, Codex, Cursor — to program AMD GPUs natively. Instead of learning ROCm's APIs through documentation, developers can ask an AI agent to write, optimize, and debug GPU code for AMD hardware.

The numbers AMD shared: ROCm.ai delivers an average 3.3x inference improvement and 2.4x training improvement compared to ROCm 7 on the same hardware. That's not a hardware gain — that's pure software optimization, which tells you how much headroom was still left on the table.

The centerpiece is Hyperloom, an open-source agentic system that automates end-to-end inference workload optimization — profiling, analysis, kernel optimization, and validation. AMD says Hyperloom can compress weeks of specialized engineering work into hours.

The partnership with Anthropic is key here. Anthropic is using Claude to accelerate ROCm development itself — AI optimizing the platform that runs AI. It's a recursive loop, and if it works, it could close the software gap with CUDA faster than anyone expects.

The CRN report includes a notable detail: AMD claims developers can migrate from CUDA to ROCm while preserving on average 75% of existing CUDA code. That's not 100%, and the last 25% is where the hard problems live, but it's a credible starting point.


Physical AI: Kria and the Robot Bet

AMD also used Advancing AI to stake its claim in physical AI — intelligence that perceives, reasons, and acts in the real world.

The Kria AI Robotics Developer Platform is the first open, turnkey platform for autonomous robotics that combines CPU, GPU, NPU, and FPGA compute on a single platform. That's important because robotics workloads span the full range: NPU for perception, GPU for reasoning, CPU for planning, FPGA for real-time control loops. Most robotics platforms force you to stitch together separate chips from different vendors.

AMD claims Kria AI solutions can handle 2.3x more concurrent AI agents than NVIDIA's Jetson T5000, with over 8,000 control decisions per second and sub-100-millisecond vision-language-action reasoning.

The new Robotics Partner Network includes World Wide Technology, Bosch Rexroth, and MulticoreWare, with no licensing or membership fees. AMD is treating robotics like an ecosystem play — get the hardware into developers' hands, make the software open, and let the community build.


What This Changes

AMD's Advancing AI 2026 presentation matters for three reasons that go beyond product specs.

First: the AI supply chain is diversifying. For two years, the industry has been warning about single-supplier risk in AI compute. NVIDIA's margins reflect its near-monopoly. AMD just demonstrated that the alternative isn't just viable — it's being adopted at gigawatt scale by the most demanding customers in the world. That changes pricing power, supply dynamics, and the speed at which AI infrastructure can scale.

Second: the architecture is catching up. CDNA 5's Wave32 shift, Helios's UALoE fabric, and ROCm.ai's agentic optimization aren't "me too" moves. They're genuine architectural innovations that address the specific needs of agentic AI workloads — low latency, high branching, variable compute patterns. AMD is no longer following NVIDIA's architectural lead. It's making its own bets.

Third: the roadmap goes to 2030. AMD committed to an annual cadence through the end of the decade. Zen 7 "Florence" CPUs in 2028. Zen 8 "Ravenna" in 2030. MI500 GPUs in 2027. MI600 in 2028. Helios 500 and 600 rack-scale refreshes on the same cadence. This isn't a one-generation push. It's a sustained, multi-generational commitment to competing at the frontier of AI infrastructure.

The $2 trillion TAM projection by 2030 is ambitious, but when you add up data center CPUs, AI accelerators, networking, edge AI, and robotics — all markets AMD is now actively targeting — the math isn't crazy.


⚠️ Limitations & Caveats

Every product launch has caveats, and AMD's deserves honest scrutiny.

  1. All performance numbers are vendor-supplied. The 34x token throughput gain vs MI355X, the 30% better tokens/dollar vs Rubin, the 245% agentic AI lead over Intel — these are all AMD's claims, not independent benchmarks. We won't know the real numbers until both Helios and Vera Rubin NVL72 are in the hands of third-party testers.

  2. CUDA's moat is deeper than benchmarks. NVIDIA's software ecosystem is 15+ years deep. ROCm has made enormous progress, and ROCm.ai is genuinely clever, but there are thousands of CUDA libraries, tools, and optimizations that don't have ROCm equivalents. The 75% code preservation claim is promising but the remaining 25% contains the edge cases that make or break production deployments.

  3. Ship dates matter. Helios starts shipping Q3 2026, ramping Q4. Vera Rubin is targeting a similar window. The company that ships first at volume gains real advantages — but AMD is also asking customers to bet on a platform that, for many, is their first AMD GPU deployment at scale. Ramp-up friction is real.

  4. The partnership commitments are impressive but non-exclusive. OpenAI is still NVIDIA's biggest customer. Meta runs NVIDIA GPUs at massive scale. Anthropic's 2GW deal is real and substantial, but it doesn't mean any of these companies are abandoning NVIDIA. They're diversifying — which is good for AMD but doesn't guarantee the kind of volume that would fundamentally reshape market share.

  5. Power density is becoming the binding constraint. A 2.9-EFLOP rack draws enormous power. As AI clusters scale to gigawatt deployments, the bottleneck shifts from chip performance to power delivery, cooling, and grid capacity. AMD's efficiency claims are important, but they need to be validated in real data center conditions.

  6. Robotics is still early. The Kria platform is impressive on paper, but the robotics market is fragmented, the software ecosystem is immature compared to CUDA for AI, and the path from "developer platform" to "production at scale" is measured in years, not quarters.


🎯 The Bottom Line

AMD didn't just launch faster chips at Advancing AI 2026. It launched a credible alternative to the NVIDIA monopoly — and convinced OpenAI, Anthropic, and Meta to sign on. The MI455X and Helios rack are competitive on raw specs. CDNA 5's Wave32 architecture is a genuine innovation for agentic AI. And the roadmap to 2030 says this isn't a one-off. Whether AMD can execute on software, ship at volume, and close the CUDA gap will determine if this is a real turning point or just the most impressive also-ran the industry has ever seen. But for the first time in years, the answer isn't obvious. And that's progress.


📚 Sources

  1. [AMD Investor Relations] — AAI 2026: AMD Delivers Full-Stack Compute for the Agentic AI Era (Official Press Release). https://ir.amd.com/news-events/press-releases/detail/1294/aai-2026-amd-delivers-full-stack-compute-for-the-agentic-ai-era

  2. [HotHardware] — AMD Announces EPYC 9006 CPUs, Instinct MI400X GPUs, And More At Advancing AI 2026. https://hothardware.com/news/amd-advancing-ai-2026-epyc-9006-instinct-mi400x-helios

  3. [CRN] — AMD Advancing AI 2026: Top News On AI Chips, CPUs, Robotics. https://www.crn.com/news/ai/2026/amd-advancing-ai-2026-top-news-on-ai-chips-cpus-robotics

  4. [GlobeNewsWire] — AAI 2026: AMD Delivers Full-Stack Compute for the Agentic AI Era. https://www.globenewswire.com/news-release/2026/07/23/3332491/0/en/AAI-2026-AMD-Delivers-Full-Stack-Compute-for-the-Agentic-AI-Era.html

  5. [Wccftech] — Watch The AMD "Advancing AI 2026" Event Live Here. https://wccftech.com/watch-amd-advancing-ai-2026-event-live-here/

  6. [CNBC] — AMD to invest up to $5 billion in Anthropic as part of computing power deal. https://www.cnbc.com/2026/07/22/amd-anthropic-ai-chip-investment.html

  7. [NVIDIA Newsroom] — NVIDIA Kicks Off the Next Generation of AI With Rubin. https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer

All claims verified against Gold-tier (AMD official press release, NVIDIA official newsroom) and Silver-tier (HotHardware, CRN, Wccftech, CNBC) sources. Each source URL was scraped and confirmed accessible. Last verified: July 24, 2026.

·