Morgan Stanley Insights: AI Price War Escalates to Compute Layer, Why Amazon and Google Are Still Bullish?

TL;DR
· According to Morgan Stanley, when 1GW of GB300 hash power is used for model API, the Model Layer ROIC is approximately 20%-60%.
· If GPUs are rented directly from a cloud platform with a 75% utilization rate, the IaaS ROIC is about 23%-39%.
· The report continues to favor Amazon and Google; Meta remains a Top Pick, but its AI returns are more driven by advertising, proprietary applications, and an open model ecosystem.
In its latest report, Morgan Stanley's Brian Nowak's team made a counterintuitive judgment: open-weight models are suppressing AI token prices, but platforms with computational power, scheduling capabilities, and enterprise access may still achieve significant returns from AI capital expenditures.
The report calculated two sets of accounts separately.
At the model API layer, Morgan Stanley based its analysis on a 1GW GB300 data center, assuming a token price per million ranging from $1 to $2.5, GPU throughput per second ranging from 2000 to 3500 tokens, and allocating 50% to 80% of the hash power for inference. Under different scenario combinations, the Model Layer ROIC is approximately 20% to 60%.
At the cloud platform layer, the report assumed a 1GW data center with around 410,000 GB300 GPUs, a GPU utilization rate of 75%, and a leasing price of $7 to $10 per hour. Based on this, the IaaS ROIC for cloud platforms renting out GPU power is approximately 23% to 39%.
These two sets of numbers together explain Morgan Stanley's key insight: as AI model prices drop, it first impacts the unit economics of model providers; however, for cloud platforms like Amazon, Google, enterprises' adoption of AI still requires GPUs, storage, databases, security, and model deployment tools. While models may become cheaper, the infrastructure that hosts the models does not lose value as a result.
The direct impact of open-weight models is to lower the barrier for enterprises to use and deploy AI.
Unlike closed models, open-weight models allow users to download model weights, fine-tune or retrain them using private data, and choose to deploy them locally, in the cloud, or through APIs. Enterprises do not have to rely solely on a single model provider and can more flexibly control data, costs, and deployment methods.

The difference between the Open Weight and Closed Weight models, including whether the weight can be downloaded, support for fine-tuning with private data, and the ability to run on proprietary hardware.
The trade-off is that the model API is becoming increasingly difficult to sustain high value. As low-cost models approaching state-of-the-art capabilities proliferate, model labs need to compete for customers through lower prices, quicker responses, and a more comprehensive toolchain. The report mentions that the latest offering from Meta, Muse Spark 1.2, and other low-cost near-state-of-the-art models are further intensifying price competition in the model layer.
This is also why the market is concerned about AI capital expenditure returns. Tech companies have invested billions of dollars in purchasing GPUs, building data centers, and locking in power supply. If the inference prices continue to drop, the ultimate return will increasingly rely on three variables: GPU utilization, single-card token throughput, and whether AI workloads can drive additional cloud service revenue.
Morgan Stanley believes that the price drop will squeeze model layer margins but does not necessarily mean that the overall ROI of AI infrastructure will deteriorate in sync. Cheaper, more flexible models may also drive more enterprises to adopt AI, thereby expanding the demand for underlying compute power and cloud services.
The ROI calculation at the model API layer revolves around three core variables: token price, GPU throughput, and the share of compute power used for inference.
Morgan Stanley uses approximately $1.75 per million tokens as the baseline price and conducts scenario analysis in the range of $1 to $2.5. The report references the public inference benchmark of open weight models such as DeepSeek-V4-Pro 1.6T on GB300 hardware, with the assumption of single GPU throughput ranging from 2000 to 3500 tokens per second.
In the baseline scenario, the report assumes 65% of the compute power is used for inference, single GPU processes 2750 tokens per second, and the price per million tokens is $1.75, corresponding to approximately $40 billion in revenue per GW. The revenue for the complete scenario range is approximately $23 billion to $52 billion per GW.
Based on this, the report estimates the model layer ROIC for 1GW of GB300 compute power to be approximately 20% to 60%.

Model API Revenue and Cost Framework. Different token prices, throughputs, and inference compute allocations correspond to revenues of approximately $23 billion to $52 billion per GW.
It is important to note that the 20% to 60% range is not automatically achieved at a fixed $1.75 price but is a scenario range formed by different price, throughput, and compute allocation combinations.
Whether this account can be realized depends on whether efficiency gains can outpace price declines. If the token price continues to plummet rapidly and GPU utilization and throughput do not increase accordingly, the model-layer returns will significantly shrink.
For cloud platforms like Amazon and Google, a more direct source of revenue is renting out computing power rather than bearing the price risk of model APIs themselves.
Morgan Stanley assumes that a 1GW GB300 data center can accommodate approximately 410,000 GB300s. At a 75% utilization rate, if the hourly rental price per GPU is $7 to $10, the corresponding annual revenue is about $18.9 billion to $27 billion per GW.
After deducting costs such as IT equipment depreciation, non-IT asset depreciation, electricity, labor, and maintenance, the report estimates that the after-tax operating profit per GW for cloud platforms is approximately $8.9 billion to $15.3 billion, corresponding to an ROIC of about 23% to 39%.

Cloud platform GB300 GPU rental model. An hourly rental price of $7 to $10, corresponding to an ROIC of approximately 23% to 39%.
This set of calculations more directly supports Morgan Stanley's optimistic view of cloud platform AI capital expenditure.
Even if model API prices decrease, enterprises still need to rent GPUs to run open-weight models. For cloud platforms, the key is not how much a single token can be sold for, but whether scarce computing power can maintain high utilization rates and rental prices.
Morgan Stanley also emphasizes that lower-priced models may bring more usage. Even if unit prices and profit margins decline, as long as the overall scale of AI workloads expands, cloud platforms' absolute EBIT may still grow.
Morgan Stanley believes that cloud platforms can withstand price wars for four main reasons.
First, computing power remains scarce. Even if companies opt for open-weight models, they still require GPUs, inference clusters, storage, networking, and a secure environment. Large enterprises can build their own infrastructure, but more customers will still rent computing power and host models through AWS, Google Cloud, or Azure.
Second, cloud platforms can improve system efficiency through scale. Larger memory capacity, faster interconnects, MoE architecture, and request batching can all increase the token throughput per GPU. Large cloud platforms have a higher volume of enterprise customer traffic, making it easier to consolidate and schedule requests, thereby increasing GPU utilization.
Third, custom chips can lower service costs. Amazon has Trainium, Google has TPU. Custom chips and software optimizations allow cloud platforms to have greater flexibility in price competition and reduce reliance on a single GPU vendor.
Fourth, AI workloads will also drive other cloud service revenues. When enterprises deploy AI, they typically also need to purchase object storage, databases, vector search, RAG, security governance, identity verification, and monitoring services. These additional services can expand the cloud revenue from each AI workload and increase overall profit contribution.
Therefore, while open weight models may turn API token revenue into a "loss leader" product, they may also lead to increased GPU rentals, data storage, and database needs.
Based on the above analysis, Morgan Stanley maintains an "overweight" rating for Amazon and Alphabet with target prices of $335 and $400, respectively.
Amazon's strength lies in its AWS enterprise customer base, Trainium custom chip, and complete cloud service ecosystem. Google, on the other hand, has both TPU, Gemini models, search and YouTube application scenarios, and Google Cloud's enterprise distribution capabilities. For both companies, the return on AI capital expenditure mainly comes from compute rentals, model hosting, and cloud service bundling.
Meta's logic is somewhat different.
Morgan Stanley still lists Meta as a Top Pick, with an "overweight" rating and a target price of $775, but Meta is not an enterprise cloud platform in the sense of AWS or Google Cloud. Its AI investment is mainly realized through an open model ecosystem, advertising efficiency, proprietary applications, and data center capabilities.

Meta's target price and risk-reward assumptions.
The report suggests that Meta's AI investment may improve user engagement, Reels monetization, and advertising measurement capabilities, providing growth opportunities for new products. However, if data center construction is poorly executed, capital intensity continues to rise, and additional compute power fails to generate sufficient returns, it could drag down free cash flow and valuation.
The report does not deny the AI price war.
As low-cost open-weight models approach state-of-the-art capabilities, enterprises will pay more attention to cost, latency, deployment flexibility, and data sovereignty. Relying solely on high-priced APIs in the Model Lab to generate high profits will become increasingly challenging, and the importance of toolchains, enterprise distribution, use cases, and ecosystems will continue to rise.
Cloud platforms also face the pressure of heavy assets. A 1GW GB300 data center involves chips, data halls, power, networking, and depreciation. Any increase in costs at any stage will impact the ultimate return. If enterprise AI application diffusion lags behind expectations, GPU utilization is insufficient, and capital expenditure becomes harder to convert into profits.
Throughput improvements will not automatically translate into gains. Hardware upgrades, MoE architecture, and request batching can all enhance efficiency but require cooperation among models, chips, networks, software stacks, and actual customer requests.
Hence, Morgan Stanley's conclusion is not "AI price war has no impact," but rather: when token prices drop, the model layer is the first to come under pressure; platforms that control scarce computing power, system optimization capabilities, and enterprise cloud on-ramps are better positioned to absorb price competition through GPU leasing, economies of scale, and additional cloud services.
What truly needs to be observed next is the throughput and utilization rate improvement on cloud platforms to see if they can continue to outpace the speed of AI price decline.
Recommended
MicroStrategy Selling Bitcoin Not Causing Price Drop, Is STRC's Rebound Really a Positive Sign?
Aug 13, 17:21
Bitwise Chief Investment Officer: Revenue is King, Crypto Valuation Logic is Being Rewritten
Aug 13, 17:19
Hyperliquid, Will It Also Support Coin Staking Rewards?
Aug 13, 16:00
Morgan Stanley Interpretation of Nebius Q2 Financial Report: AI Cloud Revenue Surges by 514%, Can the 5GW Expansion Promise Be Fulfilled?
Aug 13, 15:09
Intel CEO Pat Gelsinger's Latest Interview: After Missing Mobile, Cloud, and AI, We Can't Miss the Next Wave
Aug 13, 13:44
Why Did NeoCloud Experience the Largest Gain During This Round of Tech Stock Rebound in the U.S. Market?
Aug 13, 13:23