Morgan Stanley: AI Memory Shortage Will Persist for Years; Three Major Paths to Break the Deadlock; Structural Opportunities in Storage and Heterogeneous Computing Power
Morgan Stanley released a semiconductor industry report stating that DRAM memory shortages will persist throughout this round of the AI industry cycle, and AI computing power expansion will not wait for new wafer fabs to be built and put into production.
According to The Smart Finance APP, Morgan Stanley released a semiconductor industry research report stating that the shortage of DRAM memory will persist throughout the current AI industry cycle, and the expansion of AI computing power will not wait for new wafer fabs to be built and put into production. The industry is actively bypassing the memory bottleneck through three technological paths: hardware downscaling, inference architecture disaggregation, and CXL memory pooling, continuously advancing computing power construction under supply constraints. The firm remains optimistic about storage leaders such as Micron (MU.US) and SanDisk (SNDK.US), and also favors incremental opportunities in the CXL interconnection and heterogeneous inference tracks, recommending Astera Labs (ALAB.US), Marvell (MRVL.US), Cerebras (CBRS.US), and Nvidia (NVDA.US).
Morgan Stanley emphasized that the current AI memory shortage is not a short-term disruption, but a structural contradiction where the growth of computing performance far exceeds the speed of memory supply. The size of leading-edge large models doubles every six months, with context window annual growth rates reaching 5-6 times, combined with continuously rising concurrent inference demand, which has resulted in explosive growth in memory capacity and bandwidth requirements. However, the construction cycle for DRAM wafer fabs stretches over several years, supply flexibility is extremely low, and the shortage situation will persist for years. Nvidia CEO Jensen Huang has also publicly stated that the industry needs to shift its approach and address memory constraints through architectural innovation, rather than just waiting for capacity expansion.
In the face of tight HBM and main memory supply, the industry's most direct response is to selectively reduce memory specifications per device (de-speccing). Taking Nvidia's Rubin architecture as an example, the LPDDR5 capacity per single machine rack was reduced from the originally planned 54TB to 28TB, and the HBM capacity per GPU was lowered from 288GB to 192GB, maintaining overall system shipments by reducing memory stack height. Morgan Stanley points out that downscaling does not eliminate the memory bottleneck, but rather shifts the pressure between memory tiers: after reducing high-speed local memory, non-high-frequency data such as KV cache is offloaded to NAND storage, and cross-GPU data interaction increases, driving demand for network interconnect bandwidth. Essentially, this complements scarce high-bandwidth memory with faster network interconnects and lower-cost storage resources.
The second path to breakthrough is inference architecture disaggregation. AI inference consists of two markedly different stages: prefill and decode. Prefill is mainly about compute consumption, while decode relies heavily on memory bandwidth. Traditional architectures handle both tasks with the same accelerator, resulting in low resource utilization. The industry is quickly moving toward heterogeneous inference, splitting the two stages onto different hardware: compute-intensive prefill is handled by general-purpose GPUs, while memory bandwidth-intensive decode is processed by dedicated acceleration chips equipped with large on-chip SRAM.
Typical examples include Cerebras's wafer-scale engine and Nvidia’s Groq LPU architecture, both of which can provide much higher memory bandwidth efficiency during the decode phase than conventional GPUs. Morgan Stanley believes heterogeneous inference will become a key evolution direction for AI infrastructure, and companies specializing in the decoding segment will see clear incremental market opportunities.
The third path is CXL, which reconstructs the memory hierarchy. CXL (Compute Express Link) uses high-speed interconnection to enable memory expansion, sharing, and pooling, decoupling memory from a single processor and becoming a core technological path to overcoming memory capacity constraints. Morgan Stanley estimates that AI demand will drive the CXL and related memory accessory chip market to reach about $6 billion by 2030, far exceeding the market size for traditional CPU memory expansion. Its core value lies in building a tiered memory architecture: the highest-frequency data remains in HBM, second-high-frequency data goes into the CXL memory pool, and low-frequency cold data is pushed down to NAND, using a cost gradient to match data access frequency.
In terms of investment themes, Morgan Stanley reiterates three main directions: First, continue overweight positions in Micron and SanDisk, since downscaling is due to supply shortages rather than weak demand, and AI memory demand will trend upwards in the long term, with the supply gap steadily digesting capacity; second, position in CXL and scale-up interconnection leaders like Astera Labs and Marvell; third, focus on beneficiaries of heterogeneous inference such as Cerebras, as well as Nvidia, which is refining its heterogeneous layout via the acquisition of Groq.
Risk Warning: AI computing power demand growth may fall short of expectations; CXL technology deployment may progress slower than expected; memory capacity expansion may exceed expectations.
Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.
You may also like
France: Fiscal and Political Uncertainty Impacts Financial Sector! Credit Risk Indicators of Three Major Banks Rise, Bond Default Insurance Costs Significantly Increase
As concerns about France's fiscal situation and political climate spread to the credit market, the credit risk indicators for major French bank bonds have risen significantly.
European sovereign debt sounds the alarm, but the stock market remains resilient! French-German yield spread posts largest weekly rise in over 30 years; Deutsche Bank warns the divergence may not last
Last week, significant pressure emerged in the European sovereign bond market, but the European stock market and corporate credit market remained relatively calm, resulting in a rare divergence between different asset classes.
Here’s the Target as XRP Eyes Right-Angled Descending Broadening Trend Breakout
BNB holds near $790 as bullish derivatives meet weak ETF demand
