Research Notes

Can Merchant And Custom Silicon Really Share The Same Rack?

Research Finder

Find by Keyword

Can Merchant And Custom Silicon Really Share The Same Rack?

Annapurna becomes the first announced collaborator on NVIDIA's NVHBM memory design. The merchant silicon versus custom ASIC framing collapses as NVIDIA opens interconnect and memory design to racks where its GPUs may not sit.

8/28/2026

Key Highlights

  • AWS plans to deploy 2 million additional NVIDIA Blackwell Ultra, Rubin, and Rubin Ultra GPUs across its global infrastructure in 2027 and 2028, following an earlier commitment at GTC 2026 to add more than 1 million starting in 2026.
  • NVIDIA Vera CPU-based infrastructure is coming to AWS, positioned alongside Graviton rather than in place of it, with volumes undisclosed.
  • Amazon's Annapurna Labs will be the first announced collaborator on NVIDIA's custom high-bandwidth memory, NVHBM, which NVIDIA says is designed to deliver up to 30% greater memory bandwidth, 15% lower HBM power, and up to 25% more usable XPU compute-die area versus standard HBM4E.
  • The companies plan AI factories for the U.S. government as a separate commitment, including 100,000 GPUs on secure AWS infrastructure for workloads classified at Impact Level 6 and above.
  • Amazon EC2 G7 instances with RTX PRO 4500 Blackwell Server Edition GPUs are cited at 4.6x inference performance over G6, with GPU-accelerated indexing on Amazon OpenSearch described as up to 9x faster at roughly a quarter of the cost.

The News

AWS and NVIDIA announced a broad expansion of their collaboration on 26 August 2026, covering GPU capacity, CPUs, scale-up interconnect, custom memory, open models, data processing, and robotics. The centerpiece is a plan to deploy 2 million additional NVIDIA GPUs across AWS global infrastructure in 2027 and 2028, roughly doubling the commitment disclosed at GTC 2026. Beneath the capacity number sits a deeper technical collaboration, with Annapurna Labs named as the first partner on NVIDIA's NVHBM custom memory technology for future Trainium silicon and AWS agreeing to offer Vera CPU-based infrastructure. Full details are in the joint announcement.

Analyst Take

The announcement leads with a number designed to be quoted: 2 million additional GPUs across AWS global infrastructure in 2027 and 2028. We read the GPU count as perhaps the least interesting element of the disclosure. The more consequential material sits underneath it, in the memory subsystem and the scale-up fabric, where NVIDIA is opening controller design and interconnect IP to a hyperscaler building its own accelerators. The obvious reading is that NVIDIA secured another hyperscaler commitment at a moment when its data center business grew 117% year over year. Our read is different, and it is a read rather than a disclosure. Neither company described this as a concession, and both have argued for customer choice for years. What changed is not the rhetoric but the coupling. That coupling cuts both ways. AWS now has merchant silicon in the load path for its largest customers. NVIDIA now has design content in a rack it does not own end to end. Both positions appear rational. Neither is the headline.

What Was Announced

The capacity commitment spans Blackwell Ultra, Rubin, and Rubin Ultra, with additional Blackwell expansion for graphics and inference through EC2 G7 instances built on RTX PRO 4500 Blackwell Server Edition parts, a first-major-cloud claim AWS has carried since GTC 2026 and now supports with performance figures. The companies are also collaborating on Spectrum networking to tune network performance across large training clusters. Vera CPU-based infrastructure is framed as an additional option for agentic workloads requiring high-performance general-purpose compute adjacent to accelerators, sitting beside Graviton rather than displacing it. Volumes were not disclosed.

The architectural disclosure that matters most is NVHBM, though it is a design collaboration rather than a shipping product. NVIDIA is moving its memory controller off the XPU compute die and into the HBM base die, and says the approach is designed to deliver materially higher bandwidth and lower memory power than standard HBM4E. Annapurna Labs is the first announced collaborator, not an exclusive one, and NVIDIA has framed NVHBM as a standard implementation available from multiple memory suppliers to NVLink Fusion customers. The collaboration extends the re:Invent 2025 NVLink Fusion commitment and would begin with Trainium4. The timing gives the efficiency claims a financial rationale rather than a purely architectural one. Amazon raised 2026 cash capital expenditure guidance to roughly $220 billion citing memory costs, and NVIDIA guided its own fourth-quarter gross margin down to 71 to 72 percent under the same pressure. Both sides of this partnership are absorbing the same input shock.

The remaining elements consolidate existing work. All NVIDIA GPU-based and Trainium-based EC2 instances, including those using NVLink Fusion, are built on the AWS Nitro System and interconnected through EFA, preserving the isolation and network behavior AWS customers already depend on. Nemotron open models continue on Amazon Bedrock and SageMaker, and NVIDIA's cuDF and cuVS libraries accelerate data processing on Amazon EMR and vector indexing on Amazon OpenSearch. The federal AI factory build is listed as its own commitment rather than as a slice of the capacity number, targeting workloads at Impact Level 6 and above, a narrower and more defensible market than the headline implies.

Market Analysis

Google is no longer keeping TPU exclusively inside its own cloud. It began delivering TPU systems to customer data centers in the second quarter of 2026 and has stood up a joint venture with Blackstone to sell TPU capacity outside Google Cloud. That is additive rather than a pivot, and Google Cloud remains among the sites running NVIDIA's Vera Rubin racks. Microsoft's Maia remains Azure-only in deployment, though reporting has Microsoft pitching the next generation to external labs. The distinction that matters is the shape of the integration. Google is selling TPU as an alternative to merchant GPUs. AWS is treating NVIDIA GPUs, Trainium, Graviton, and Nitro as one heterogeneous fleet where the parts are meant to interoperate inside a common rack rather than compete for it. That aligns with what enterprise buyers tell us they want. HyperFRAME Lens research, from our 1H 2026 State of Enterprise Infrastructure and Operations survey of roughly 520 qualified I&O leaders, finds 49% prioritize platform interoperability over raw AI acceleration.

For NVIDIA, we think the strategic gain is less about GPU units than about position, though this is our inference and not a disclosed arrangement. No commercial terms were announced. If NVLink Fusion and NVHBM become standard furniture in semi-custom racks, NVIDIA becomes a landlord in buildings it does not own, present in the architecture whether or not it is present in the accelerator. Annapurna is not the only route in either. NVIDIA's multibillion-dollar investment in Marvell, one of the two dominant custom ASIC design houses and a long-standing partner on Amazon's accelerator programs, opens a second entry point into custom racks NVIDIA does not architect.

There is a reasonable objection. Sharing controller design and reclaimed die area could make Trainium4 a stronger competitor to Rubin using NVIDIA's own engineering. We think that understates the alternative, which is not a world where Trainium stalls but one where it ships on commodity memory inside a rack NVIDIA never touches.

The demand backdrop supports both fleets expanding at once, which makes the reading that Trainium is underperforming hard to sustain, though that remains an inference. Anthropic and OpenAI have each committed to multi-year, multi-gigawatt Trainium capacity, anchoring AWS custom silicon demand independent of anything announced here, and AWS reported roughly $496 billion in contracted backlog at the close of the second quarter while management says capacity remains short of demand. Nothing in this week's release is a Trainium volume or utilization figure. When the binding constraint is power, packaging, and memory rather than architectural preference, buying more of everything is the coherent response.

Looking Ahead

The key trend we'll be monitoring is whether NVHBM becomes a pattern or stays an exception. Annapurna Labs is first but explicitly not exclusive, and whether other XPU designers take up the standard implementation will indicate how much of the semi-custom rack NVIDIA ends up defining. Trainium4 silicon built around NVIDIA's controller design would be the clearest test yet of whether merchant and custom accelerators can share a rack without one cannibalizing the other. Physical AI is worth tracking as a demand category in its own right rather than as a subsidiary of the agentic framing. Amazon Robotics has announced adoption of the Jetson, Omniverse, and Isaac stack across simulation, synthetic data generation, and real-to-sim validation. That is a stack commitment from a named fleet, not yet a documented production rollout, and the gap between those two is the thing to watch. NVIDIA's Q3 FY2027 guide of approximately $108 billion sets a demanding bar that these commitments will need to help clear.

Author Information

Stephen Sopko | Analyst-in-Residence – Semiconductors & Deep Tech

Stephen Sopko is an Analyst-in-Residence specializing in semiconductors and the deep technologies powering today’s innovation ecosystem. With decades of executive experience spanning Fortune 100, government, and startups, he provides actionable insights by connecting market trends and cutting-edge technologies to business outcomes.

Stephen’s expertise in analyzing the entire buyer’s journey, from technology acquisition to implementation, was refined during his tenure as co-founder and COO of Palisade Compliance, where he helped Fortune 500 clients optimize technology investments. His ability to identify opportunities at the intersection of semiconductors, emerging technologies, and enterprise needs makes him a sought-after advisor to stakeholders navigating complex decisions.