AI Memory Constraints Drive Tiered Storage and CXL Designs
Vera Rubin systems pair HBM4 with LPDDR5X, storage and high-speed links as longer AI context windows increase memory requirements.

AI infrastructure is moving toward tiered memory and storage designs as larger AI context windows increase pressure on high-bandwidth memory and force data across multiple system layers.
The Vera Rubin NVL72 rack combines 72 Rubin GPUs and 36 Vera CPUs with 20.7TB of HBM4 and 54TB of LPDDR5X. The HBM4 total works out to about 287.5GB per GPU, broadly matching a 288GB configuration.
HBM sits close to AI accelerators and provides high bandwidth for demanding workloads. System memory, flash storage and shared-memory systems handle data that cannot remain in local HBM, creating a broader architecture for managing capacity and data movement.
Longer context windows increase the amount of key-value cache that AI systems must retain or retrieve during inference. NVIDIA’s BlueField-4 STX storage architecture is designed to store and retrieve large-scale AI key-value cache data, extending available capacity beyond GPU HBM.
The shift could increase demand for storage, memory and connectivity components as data-center operators distribute workloads across more tiers. Compute Express Link, or CXL, supports memory expansion, sharing and pooling, allowing data centers to allocate memory across CPUs, GPUs and other accelerators.
Marvell’s Structera S CXL switch is designed for rack-level memory pooling and dynamic allocation. The company expects its 30260 switch to begin customer sampling in calendar third quarter of 2026. Astera Labs’ Leo X-Series targets fabric-attached memory for AI inference and key-value cache offload. CXL memory adds capacity and flexibility but does not match the bandwidth of local HBM.
High-speed interconnects remain central as workloads become more distributed. NVLink 6 provides 3.6TB/s of bidirectional GPU-to-GPU bandwidth per GPU, supporting communication among accelerators in large AI systems.
AMD and Cerebras are also separating inference workloads. AMD Helios handles prompt processing and large context windows, while Cerebras’ Wafer-Scale Engine manages memory-bandwidth-intensive decoding and token generation. The system is planned for availability through Cerebras Cloud in the second half of 2026 and is designed to deliver up to five times higher tokens per second per watt.
“AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach,” AMD Chair and CEO Lisa Su said.
NVIDIA founder and CEO Jensen Huang described Vera Rubin as “a generational leap — seven breakthrough chips, five racks, one giant supercomputer — built to power every phase of AI.”