Morgan Stanley Analysis: Why Is Nvidia Reducing Memory Despite AI Memory Shortages?
TL;DR: AI memory shortages may persist throughout the cycle. Morgan Stanley believes that the continuous growth in model size, context length, and inference concurrency will keep absorbing new storage supply. Nvidia has started offering "reduced configuration" options for Rubin. Some HBM and LPDDR5 capacities may be lowered, but demand has not disappeared; instead, it is shifting towards NAND, DRAM, and high-speed interconnects. AI inference architecture is being restructured. Separating Prefill and Decode is expected to improve hardware utilization and create opportunities for technologies like Cerebras and Nvidia Groq. CXL is expected to become a new growth point. Morgan Stanley predicts that by 2030, the related semiconductor market will reach approximately $6 billion, with Astera Labs and Marvell as potential beneficiaries. The key for storage stocks is the duration of the cycle, not the short-term price increases. Morgan Stanley maintains an overweight rating on Micron and SanDisk, with the main risk being a slowdown in AI investment and data center construction.
Editor’s Note: The next bottleneck for AI may not be the number of GPUs, but memory. As the scale of large models expands, context windows grow, and AI agents continue to consume more inference resources, the supply pressure on HBM, DRAM, and NAND is being transmitted throughout the data center. More importantly, in the face of storage shortages, Nvidia is exploring lower memory configuration product solutions for the next-generation Rubin platform to maintain system delivery and deployment pace.
This raises a seemingly contradictory question: If the demand for storage in AI is so strong, why are chip manufacturers starting to consider reducing memory capacity? Does memory "reduction" mean that demand is about to peak, or does it indicate that supply bottlenecks have become severe enough to force the entire industry to adjust its technological direction?
Morgan Stanley semiconductor analysts Joseph Moore and others published a report on October 5 titled "How Can the AI Ecosystem Work Around Memory Bottlenecks?" They suggested that storage shortages may persist throughout the current AI cycle, but AI construction will not wait for DRAM wafer fabs to complete expansion. The industry needs to find ways to bypass existing supply constraints by reducing some product configurations, splitting inference tasks, and improving memory sharing efficiency.
This means that competition in AI hardware may shift from simply increasing GPU and memory capacity to improving the resource utilization efficiency of the entire computing system. Storage shortages do not necessarily weaken the long-term demand for storage manufacturers but may change the value distribution along the supply chain. In addition to storage companies like Micron and SanDisk, companies like Cerebras, Astera Labs, and Marvell, which have specialized computing architectures and interconnect technologies, may also gain new growth opportunities from this.
The following is the original text compilation:
AI infrastructure construction is facing an increasingly difficult problem: computing power can be expanded by purchasing more GPUs, but memory supply cannot increase at the same speed.
Morgan Stanley believes that AI demand may continue to absorb a large amount of new storage supply in the foreseeable future. Even if the degree of shortage fluctuates at different times, the expansion of DRAM capacity will still struggle to quickly eliminate the supply-demand gap.
In discussions with companies in the computing industry, Morgan Stanley found that how to bypass memory bottlenecks has become an increasingly important topic. Nvidia CEO Jensen Huang also mentioned the necessity of alleviating supply constraints through the collaborative design of computing, networking, and storage systems.
Analysts believe this is not a warning of weakening storage demand but a reality that the entire AI industry must face: when memory cannot be supplied as originally planned, companies must find new system architectures to allow existing hardware to continue supporting AI growth.
1. Nvidia Starts to Reduce Memory for Rubin, but Storage Bottlenecks Remain
The most direct approach is to reduce the memory capacity used by each server or GPU.
Morgan Stanley points out that the tight supply and rising prices of DRAM and NAND are forcing the AI industry chain to reconsider product configurations. For chip manufacturers, rather than sticking to original specifications and causing systems to be unable to be delivered as planned, it is better to offer versions with different memory capacities, allowing customers to choose based on actual needs.
The report predicts that Nvidia may offer multiple lower memory configurations for the Rubin platform.
In terms of rack-level memory, the originally planned LPDDR5 capacity for Rubin was about 54TB, while some new configurations may drop to 28TB, corresponding to SOCAMM2 memory module capacity decreasing from 192GB to 96GB.
For HBM, the original expectation for Rubin's single GPU was to be equipped with 288GB HBM4, using an 8-group, 12-layer stacked design; Morgan Stanley expects Nvidia may add a product version equipped with 192GB HBM4, reducing the number of stacked layers to 8.
For Rubin Ultra, the report suggests that after adjusting to a dual compute chip solution, 512GB is a more appropriate comparison benchmark, with future options possibly ranging from about 192GB to 384GB.
These are all judgments made by Morgan Stanley as of the report's publication date, and not all specifications have been officially confirmed by Nvidia.
From a cost perspective, this strategy has its rationale. Reducing the number of HBM stacked layers can decrease capacity, and when conditions such as the number of interfaces and pin rates remain unchanged, theoretical bandwidth does not necessarily decline simultaneously.
However, capacity and bandwidth address two different issues. For applications with smaller models and shorter contexts, reducing capacity may have limited impact; but as models grow larger and inference tasks become more complex, insufficient memory capacity may still lead to more data needing to be transferred between different storage levels, increasing latency or GPU usage.
More importantly, reducing HBM does not make the data that AI originally needed to process disappear; it only forces this data to find new storage locations.
For example, when the HBM capacity of a GPU is insufficient, some data may need to be stored in main memory; when rack-level LPDDR5 capacity decreases, some KV Cache (key-value cache) demand may shift to NAND flash storage.
KV Cache is a cache used to store historical computation information during the inference process of large models. As context lengthens and concurrent requests increase, its capacity demand also rises.
Storage demand does not disappear, but shifts to other levels.
Morgan Stanley points out that while Nvidia reduces some LPDDR5 configurations, the industry chain has already observed new demand from NAND. The reduction in HBM capacity may also increase data transfer between GPUs, putting greater pressure on high-speed interconnect networks.
This means that Nvidia's reduction strategy may temporarily alleviate DRAM supply constraints but simultaneously increase the demand for NAND, interconnect chips, and larger GPU clusters.
This pressure transfer has its long-term background.
According to historical data from Epoch AI cited in the report, the parameter scale of cutting-edge models in the era of large language models has shown a trend of doubling approximately every six months. Context windows have also experienced rapid expansion. At the same time, more AI applications are shifting from simple Q&A to programming, complex reasoning, and long-running agent tasks, further increasing memory usage.
Morgan Stanley believes that model size, context length, and inference concurrency will continue to drive storage demand growth. Therefore, reducing some product configurations is more likely a transitional solution during periods of supply tightness rather than a signal of a long-term decline in AI memory demand.
2. AI Inference Begins to "Decouple" Tasks, GPUs No Longer Handle All Tasks
If reducing memory configurations is a short-term response, then changing the computation method of AI inference may bring deeper industry changes.
Morgan Stanley's second focus is on Disaggregated Inference, which involves decoupling different inference tasks that were originally concentrated in the same computing system and assigning them to more suitable hardware for execution.
Traditional large model inference mainly consists of two stages.
The first stage is Prefill, which processes user input prompts, files, or historical context and performs a large amount of parallel computation. This stage typically relies more on computational performance.
The second stage is Decode, where the model generates output tokens one by one. This process requires repeated access to model weights and KV Cache, thus relying more on memory bandwidth, capacity, and data access latency.
In the past, the same GPU often handled both types of tasks simultaneously. However, as inference demands expand, this configuration may not always be the most efficient choice.
Morgan Stanley believes a more reasonable direction may be to have compute-intensive hardware handle Prefill, while architectures optimized for low latency and high bandwidth take on Decode.
This approach can not only improve hardware utilization but also allow data centers to expand the two types of computing resources separately, without needing to configure identical GPUs and HBM for all tasks.
Technical distinctions between Prefill and Decode
One of the most direct potential beneficiaries is Cerebras.
Cerebras uses a wafer-scale processor architecture, integrating a large number of computing units and SRAM on the same chip. SRAM (Static Random Access Memory) typically has lower capacity density than DRAM but features low latency and high bandwidth, making it suitable for certain inference tasks that require frequent data access.
By placing computing units closer to the data, Cerebras can reduce the need for external memory access and improve processing speed for specific inference stages.
Morgan Stanley points out that Cerebras has collaborated with AMD and AWS to explore combining different processors into a unified inference system.
In the joint solution between AMD and Cerebras, AMD Helios is responsible for the more compute-intensive Prefill and large context processing, while the Cerebras wafer-scale engine handles Decode. The report predicts that this solution will enter production in the fourth quarter of 2026.
AWS adopts a similar approach, processing Prefill with Trainium and then handling Decode with Cerebras, with related combination solutions expected to enter Amazon Bedrock in the first quarter of 2027.
According to specific performance data disclosed by Cerebras and AMD, the combination is expected to achieve up to five times the throughput improvement while maintaining Cerebras' inference speed. This does not mean that all inference tasks will receive the same level of improvement, but it indicates that configuring hardware for different tasks may significantly enhance system economics.
For Cerebras, this change not only means more opportunities for chip sales. Since the company also operates its own inference cloud service, if the same hardware can generate more tokens, the cost per token may decrease, thereby improving gross margins or providing room for lower customer prices.
NVIDIA is also exploring similar directions.
By integrating Groq's relevant technologies, NVIDIA is combining HBM-based GPUs with LPU architectures that use high-speed SRAM. In the proposed solution described in the research report, the Rubin GPU is responsible for pre-filling and some decoding calculations that require large cache capacity, while the Groq architecture handles computations that are more suited to its low-latency characteristics, coordinated by NVIDIA Dynamo software across different processors.
It is important to note that the agreement reached between NVIDIA and Groq in 2025 is for technology licensing and talent acquisition arrangements, not a complete acquisition of Groq.
Morgan Stanley believes that as inference spending grows faster than training, heterogeneous computing architectures optimized for different inference stages may gain larger market space. According to the research report's infrastructure spending forecast, by 2030, inference is expected to account for about 56% of AI infrastructure spending, while training will account for about 44%.
This also brings a new competitive logic to the AI chip industry: in the future, the metrics for measuring the competitiveness of an AI system may not only include peak computing power but also how many tokens are generated per second and the cost of generating each token.
III. CXL Market Expected to Reach $6 Billion, Memory Interconnect Becomes New Opportunity
In addition to changing the allocation of inference tasks, Morgan Stanley also focuses on how to improve the efficiency of existing memory resources. This is precisely where CXL technology has its opportunity.
CXL (Compute Express Link) is a high-speed interconnect protocol that allows processors to access external memory more flexibly without always relying on directly connected local DRAM.
In traditional servers, memory resources are often bound to specific CPUs. Even if a server has a large amount of idle memory, it is difficult to directly supply it to another server. CXL improves the overall utilization of DRAM in data centers through memory expansion, sharing, and pooling. Among these, memory expansion can provide additional capacity for a single processor; memory sharing allows multiple processors to access the same memory resources; and memory pooling organizes dispersed memory into a unified resource pool, dynamically allocating based on workload.
This technology has been more commonly applied in traditional CPU computing environments, as CPU workloads are relatively more tolerant of the additional access latency brought by external memory. In contrast, AI GPUs have long relied on HBM to provide extremely high data bandwidth, making it difficult for externally connected DRAM via CXL to directly replace HBM.
However, as the demand for AI inference changes, the value of this architecture is rising.
For example, long-context inference requires storing a large amount of KV Cache, but not all cached data needs to remain in the fastest HBM. The system can keep performance-sensitive data in HBM while transferring some lower-access-frequency data with relatively relaxed latency requirements to external DRAM.
This can alleviate the capacity pressure on expensive HBM while improving the utilization of existing DRAM.
Morgan Stanley believes that AI is driving CXL from traditional server memory expansion to a broader market for accelerator memory connections and sharing.
In the past, Astera Labs estimated the potential market size for CXL memory controllers to be over $4 billion. Morgan Stanley currently predicts that as demand for AI inference, KV Cache offloading, and rack-level memory pooling grows, the potential market size for CXL and related memory connection semiconductors could reach about $6 billion by 2030.
It is important to emphasize that this figure is an analyst's forecast of potential market space, not actual market revenue.
In terms of specific companies, the research report focuses on Astera Labs (ALAB) and Marvell (MRVL).
Astera Labs' opportunities mainly come from its Leo memory controller products. Morgan Stanley points out that the company expects standardized and customized Leo products to begin scaling in 2027 with two large U.S. cloud service provider customers, including designs for KV Cache offloading aimed at AI inference.
Marvell, on the other hand, is laying out memory expansion and rack-level memory pooling through Structera X and Structera S, respectively. The company previously expected that its CXL business could contribute over $1 billion in revenue around 2028.
For these two companies, the AI storage shortage may create new commercial demand for CXL technology, which has previously been slow to promote.
However, Morgan Stanley also acknowledges that the actual scale of the CXL market remains difficult to predict accurately. Different cloud vendors may choose non-CXL interconnect technologies, larger capacity HBM, or flash-based caching solutions, and the actual market adoption speed may differ significantly from current expectations.
Therefore, whether the investment logic for CXL can be realized still requires observation of the actual deployment progress, product shipment volumes, and related business revenues of cloud vendors.
-- Price
IV. Memory Downscaling May Not Be Bad News for Micron, Real Risk is Slowdown in AI Construction
For the storage industry chain, NVIDIA's reduction of some memory configurations does not seem to be good news. If the HBM capacity of a single GPU decreases and the usage of LPDDR5 per rack decreases, the amount of memory that storage vendors can sell may be lower than originally planned.
Morgan Stanley acknowledges that from the perspective of maximizing short-term profits, this may indeed weaken some demand and pricing opportunities that storage vendors could have obtained. However, analysts believe this should not be simply understood as the end of the storage cycle. The key is that the current downscaling is mainly driven by supply constraints, not because customers no longer need as much memory.
When AI vendors wish to procure more DRAM but cannot obtain sufficient supply, reducing configurations can help them maintain system delivery. After supply increases in the future, vendors will still have the motivation to raise configuration levels again. This means that the currently unmet demand may create some follow-up procurement space.
Based on this judgment, Morgan Stanley is more focused on the duration of the storage industry's prosperity cycle rather than the maximum extent of short-term price increases. If AI models continue to expand, Agent applications continue to increase, and inference concurrency continues to grow, then even if DRAM manufacturers increase production capacity, the new supply may be quickly absorbed.
Therefore, Morgan Stanley maintains an Overweight rating on Micron (MU) and SanDisk (SNDK), believing that storage supply tightness will not end quickly.
However, this does not mean that the storage industry can completely escape cyclical risks. The research report believes that the most critical risk still comes from AI demand itself. If the growth rate of model development, inference applications, or computing power investments declines significantly, both the computing and storage industry chains may be impacted.
In addition, there is a more complex situation: AI demand remains strong, but data center construction is delayed due to land, power, or infrastructure constraints. In this scenario, storage products may have already been produced but cannot be digested as expected due to delays in server deployment, leading to temporary supply-demand imbalances.
Morgan Stanley believes that these construction bottlenecks may create additional disturbances for storage vendors, even causing their short-term performance to diverge from some computing chip companies. Therefore, future assessments of whether the storage cycle can continue should not only observe DRAM prices and HBM orders but also track the actual deployment progress of AI infrastructure.
The ultimate impact of the AI storage shortage may not be to reduce the use of memory in the industry but to force the industry to reconsider: which data must remain in HBM, which can be transferred to DRAM or NAND, and which computing tasks should be assigned to different types of chips.
For investors, the core variables to focus on are gradually becoming clear: whether Rubin's actual shipment configuration is reduced, whether decoupled inference can achieve commercial-scale deployment, whether CXL products can scale as planned in 2027, and whether AI data center construction continues to digest new storage supply.
If these technological adjustments can maintain AI system deployment while long-term storage demand continues to grow, then the storage and interconnect industry chains may benefit together; but if AI investment slows down, or if land and power constraints continue to hinder data center establishment, the current supply-demand tightness may face re-pricing.
This is precisely Morgan Stanley's most important judgment on this round of the AI storage cycle: the industry truly needs to solve not how to wait for more memory, but how to continue expanding AI's computing power in a context of persistent memory scarcity.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Cedears: Record Rates, Euphoria for AI, and Brazil Reshaping the Stock Map—What Could Happen Next?
WEEX TOKEN2049 Gifts: A Preview of the Limited-Edition Merch
Preview the WEEX TOKEN2049 Singapore merchandise experience, what visitors should know before collecting gifts, and where to find the latest official details.

Standard Chartered plans institutional crypto custody service in Singapore

Mr&强 Shares WEEX TOKEN2049 Event Details

Mr&强 Analyzes BTC and ETH Opening Strategies

The Killing Line of China's Large Models Has Been Cut

Wall Street Morning Report: AI Bull Market Faces Interest Rate Judgment, Funds Shift to Defensive Assets, 30-Year Treasury Bonds Become Next Test

Cardano Foundation spins out digital identity firm Veridian

UK Names Six Lead Banks for First Digital Native Gilt Pilot 'DIGIT'

Tom Lee: S&P 500 Valuation Declines, Q3 Earnings Growth Approaching 30%

BofA Says Bonds Are Attractive for the First Time, U.S. Stocks May Have Lower Returns Than Treasuries Over Next Decade

Why Asia Fell Despite Records on Wall Street

Elon Musk Net Worth Trillion Dollars: Will He Be a Trillionaire Again?

What Does a Rising VIX Mean for the Stock Market? A 2026 Investor Guide

Ondo Finance Tokenizes Pre-IPO AI Stocks: Private Market Open 24/7

Expansion of 'Blockchain Bonds' into Domestic Financial Sector: Will the Funding Paradigm Change? – Bitplanet

SpaceX AI + Space Boom: Can SPCX Rally Toward $230?

Midterm Elections Approaching: What Really Matters in the Stock Market

ICC Russia Crypto & Trade Finance Conference on October 20 in Moscow

What if the biggest Bitcoin Holders were AIs?

Singapore FinTech Festival: Digital Euro and Stablecoins in the Spotlight

Hyperliquid Integrates with Bloomberg Terminal, Receives First Payment of $14.58 Million USDC

Sun Yuchen Shares Insights on Crypto Industry Trends, Emphasizing the Integration of AI and Digital Finance

Ondo Finance Launches Tokenization Market for Private Companies

Crypto: Jay Clayton, the Man Who Took on Ripple, Named AI Tsar by Trump

Transparency Act Stalls, But Bankers Continue Crypto Deals

AI Investment Also Favors the 'Winner Takes All'... Top 1% of Companies Spend 600 Times More Than Median Companies

Nasdaq Rises 1.05% to New High, 10-Year U.S. Treasury Yield Rises to 5.34%

Cryptocurrency Shifts to a New Task: Not Just Launching Products, but Retaining Users

Wall Street Bets on "Cross Strategies" as Market Gains Volatility
Cedears: Record Rates, Euphoria for AI, and Brazil Reshaping the Stock Map—What Could Happen Next?
WEEX TOKEN2049 Gifts: A Preview of the Limited-Edition Merch
Preview the WEEX TOKEN2049 Singapore merchandise experience, what visitors should know before collecting gifts, and where to find the latest official details.










