Research Finder
Find by Keyword
VAST Data and AMD Extend Enterprise AI Inference Beyond GPU Memory
VAST's platform and AMD's compute, networking, and software portfolio combine persistent KV cache with accelerated inference to reduce recomputation in long-context workloads.
7/27/2026
Key Highlights
- VAST Data and AMD announced an expanded collaboration integrating the VAST AI Operating System with AMD Instinct GPUs, sixth-generation EPYC processors, ROCm, Infinity Context, and Pensando Pollara 400 AI NICs.
- The architecture preserves computed KV cache on VAST's distributed platform, allowing GPUs to skip prefill computation on previously processed context.
- VAST reported up to 9x TTFT improvement and up to 9.7x token throughput on an AMD Instinct MI355X system, measured against full recomputation. VAST produced the results; AMD reviewed them without independent verification.
- VAST applies native data lifecycle policies to KV cache, automating expiration and deletion of cached data containing sensitive information.
- HyperFRAME believes the collaboration strengthens AMD's enterprise inference platform by combining accelerated computing with enterprise-scale context management delivered through the VAST AI Operating System.
The News
At AMD Advancing AI 2026, VAST Data and AMD expanded their collaboration to integrate AMD compute, networking, and ROCm with the VAST AI Operating System for large-scale inference. The architecture preserves persistent KV cache beyond GPU memory through VAST's Disaggregated Shared Everything (DASE) architecture, allowing enterprise data services to participate directly in inference execution. For more information, read the VAST-AMD partnership press release.
Analyst Take
This announcement advances VAST's evolution from a high-performance data platform to an AI Operating System. Existing capabilities built on the DASE architecture, including Context Memory, unified data services, metadata management, streaming, databases, and AI orchestration, now extend directly to persistent inference context.
LLM inference depends on KV cache, which preserves computed attention state between requests. VAST reports that a 128,000-token context window generates 20 GB to 50 GB of KV cache per user depending on model and data type. Aggregate requirements across concurrent sessions exceed available GPU memory, forcing recomputation of previously processed context.
The benchmark tested GPT-OSS 120B at 120,000-token input length on a single MI355X with tensor parallelism of one, running vLLM 0.22.1, LMCache, and ROCm, connected over 800 Gbps NFSv3 with RDMA and multipathing. VAST reported up to 9x TTFT improvement and 9.7x throughput against a baseline requiring recomputation at each decode turn.
VAST tested three configurations: GPU with LMCache offloading to VAST storage, GPU with LMCache offloading to local host RAM, and GPU without cache. Local host RAM delivered the lowest P99 TTFT at every concurrency level above 20 requests, reaching 204 seconds at 300 concurrent requests against 1,918 seconds for VAST and 1,977 seconds for the no-cache baseline. The reported 9x advantage occurs at lower concurrency, where the performance curves diverge. At higher concurrency, VAST closely tracks the no-cache baseline.
Local RAM offload is bounded by host memory; the tested LMCache configuration capped host cache at 300 GB and served a single node. DASE provides a shared, persistent context tier accessible across a global namespace, with capacity independent of host memory. Any compute node draws from a common context pool without duplicating context on individual hosts. That distinction defines the deployment case for external KV cache.
Native lifecycle management reinforces VAST's architectural differentiation. Cached context contains user data subject to retention and deletion requirements. VAST applies existing lifecycle policies to expire and delete KV cache, extending metadata services, security controls, multi-tenancy, and namespace management to inference state. Enterprises deploying persistent cache inherit compliance mechanisms already governing their data.
In our view, the collaboration strengthens AMD's enterprise inference position. NVIDIA supplies comparable cache tiering natively through Dynamo and NIXL within a CUDA ecosystem. AMD closes that capability through partnership. The supporting cloud roster, including Core42, Crusoe, Vultr, TensorWave, and 5C, indicates deployment traction. TensorWave describes itself as an AMD-centric AI cloud. Vultr is notable because it positions the VAST-AMD architecture as part of a validated enterprise AI stack spanning infrastructure, accelerated compute, enterprise data, cloud operations, and AI software. Vultr runs a deliberate two-lane GPU strategy, with AMD Instinct and NVIDIA deployments within the same cloud platform. That dual footing positions Vultr to provide practical operational comparisons between AMD-native and NVIDIA-based inference architectures, rather than relying on experience with a single GPU ecosystem.
What Was Announced
VAST and AMD announced an expanded engineering collaboration integrating AMD accelerated computing with the VAST AI Operating System for production inference. Current validation centers on AMD Instinct MI355X accelerators. The software environment includes AMD ROCm, vLLM, and LMCache, maintaining persistent KV cache outside GPU memory while keeping it accessible during inference execution. VAST identified persistent KV cache as the primary constraint for long-context inference, reporting 20–50 GB of cache per 128,000-token session depending on model and precision.
VAST evaluated three approaches: persistent KV cache on the VAST platform, host-memory offload through LMCache, and full recomputation. Host memory produced the lowest latency within the limits of a single server. Built on DASE, the platform addresses a different problem by providing shared, persistent context beyond host memory and across compute nodes through a global namespace.
VAST selected sixth-generation AMD EPYC processors, codenamed Venice, for sixth-generation CBox and third-generation EBox platforms supporting DataStore, DataBase, and DataEngine services. VAST cites PCIe Gen 6 support delivering twice the generational I/O bandwidth and lower latency for database, data warehouse, and event streaming services.
AMD Pensando Pollara 400 AI NICs provide the high-speed data path connecting Instinct GPUs to the VAST platform, using NFS over TCP and NFS over RDMA to move data from GPU memory to the storage cluster.
VAST, AMD, and DriveNets are developing an AI infrastructure reference architecture built on AMD Helios rack-scale systems with DriveNets AI Fabric networking, documenting sizing guidance for training, inference, reinforcement learning, and KV cache workloads. VAST expanded software ecosystem collaboration with TensorMesh and EmbeddedLLM for production inference deployment. Native VAST data lifecycle policies automatically expire and delete KV cache data, addressing retention requirements for cached content containing sensitive or personal information.
Looking Ahead
Persistent KV cache converts inference state into managed enterprise data. Context carries retention obligations, tenancy boundaries, and access controls identical to the datasets it derives from. VAST holds an advantage here because those services exist in the platform already. Competitors building cache tiers as point capabilities will rebuild lifecycle management, multi-tenancy, and namespace controls that VAST inherits. That gap widens as regulated industries deploy agentic workloads at scale.
The benchmark does not quantify the cost tradeoff between persistent storage capacity and reclaimed GPU compute. Recomputation consumes accelerator cycles that produce no tokens. Persistent KV cache introduces storage and networking costs in exchange for higher GPU utilization. Cost per generated token under sustained inference will ultimately determine the architectural value of persistent cache.
Venice production systems will validate the CPU foundation for VAST's next-generation CBox and EBox platforms. Helios deployments will demonstrate whether the reference architecture scales beyond laboratory testing. AMD's software ecosystem will determine whether partner-supplied context management materially narrows NVIDIA's advantage in distributed inference. It will be important to see results at 1 million tokens, where VAST projects greater gains and where local RAM offload becomes infeasible.
Enterprise customers should evaluate not only accelerator performance, but also how vendors manage inference state throughout its lifecycle. Persistent context becomes infrastructure that must be governed, secured, retained, shared, and deleted alongside the enterprise data from which it is derived. Those capabilities will influence operational efficiency, regulatory compliance, infrastructure utilization, and the total cost of delivering AI services at production scale.
Don Gentile | Analyst-in-Residence -- Storage & Data Resiliency
Don Gentile brings three decades of experience turning complex enterprise technologies into clear, differentiated narratives that drive competitive relevance and market leadership. He has helped shape iconic infrastructure platforms including IBM z16 and z17 mainframes, HPE ProLiant servers, and HPE GreenLake — guiding strategies that connect technology innovation with customer needs and fast-moving market dynamics.
His current focus spans flash storage, storage area networking, hyperconverged infrastructure (HCI), software-defined storage (SDS), hybrid cloud storage, Ceph/open source, cyber resiliency, and emerging models for integrating AI workloads across storage and compute. By applying deep knowledge of infrastructure technologies with proven skills in positioning, content strategy, and thought leadership, Don helps vendors sharpen their story, differentiate their offerings, and achieve stronger competitive standing across business, media, and technical audiences.
Stephen Sopko | Analyst-in-Residence – Semiconductors & Deep Tech
Stephen Sopko is an Analyst-in-Residence specializing in semiconductors and the deep technologies powering today’s innovation ecosystem. With decades of executive experience spanning Fortune 100, government, and startups, he provides actionable insights by connecting market trends and cutting-edge technologies to business outcomes.
Stephen’s expertise in analyzing the entire buyer’s journey, from technology acquisition to implementation, was refined during his tenure as co-founder and COO of Palisade Compliance, where he helped Fortune 500 clients optimize technology investments. His ability to identify opportunities at the intersection of semiconductors, emerging technologies, and enterprise needs makes him a sought-after advisor to stakeholders navigating complex decisions.



















