← Back to news
Archived · Published 14 August 2026
The AI Bottleneck Quietly Moved From Compute to Memory
The shorthand for AI hardware scarcity has been the accelerator — the specialised processor that performs the matrix arithmetic underneath model training and inference. That framing was accurate during the period when fabrication capacity for those processors was the binding constraint. It has become progressively less accurate as a different component moved into the critical path: high-bandwidth memory, the stacked DRAM packaged directly alongside the processor that determines how much of a model can sit close enough to the compute to be useful.
The reason memory matters more than raw arithmetic throughput for serving large models is that inference is overwhelmingly memory-bound rather than compute-bound. Generating each output token requires reading the model's parameters, and the processor spends much of its time waiting for those reads rather than saturated with arithmetic. A machine with abundant compute and insufficient memory bandwidth runs at a fraction of its theoretical capability — which is why memory capacity and bandwidth per package, not headline arithmetic figures, increasingly determine what a given system can actually serve and at what cost per request.
The supply picture reflects that reassessment. Memory manufacturers have committed capital at a scale historically reserved for logic fabrication, with one major manufacturer announcing a multi-year, multi-hundred-billion commitment to AI-oriented memory capacity across Korean and American sites. The stacking and packaging processes that produce high-bandwidth memory have lower yields and longer cycle times than conventional DRAM, so capacity added is not capacity available: the lag between announcement and shipping product runs to years, which is precisely why the announcements are being made at this scale now.
The consequence for anyone buying rather than building AI infrastructure is that pricing and availability signals have become harder to read from accelerator supply alone. A period of comfortable processor availability paired with tight memory supply looks, from the outside, like an easing market while the actual serving cost per token stays flat or rises. Reading the memory market has become part of reading the AI cost curve, and it is the half that gets far less attention than it now deserves.
Defici Editorial · Tech News
This article was generated by Defici's AI editorial system.