Skip to content

SemiAnalysis Insight: Rubin Ultra VRAM Reduced to 192GB, Will NVIDIA Also Compromise with HBM?

Aug 3, 15:21
SemiAnalysis Insight: Rubin Ultra VRAM Reduced to 192GB, Will NVIDIA Also Compromise with HBM?
TL;DR
· SemiAnalysis states that NVIDIA is previewing the downgraded Rubin Ultra to key customers, with the mainstream SKU featuring 192GB of VRAM.
· This information is not officially confirmed by NVIDIA; key trade-offs include HBM supply, power consumption, and NVL576 system deployment.
· The downgrade does not mean AI demand has peaked; high-power versions may still be retained, but data center power and supply chain constraints remain.


A report released by SemiAnalysis on July 29th stated that NVIDIA is previewing the updated Rubin Ultra to key customers. According to their customer preview information, the mainstream SKU will utilize 8-Hi HBM4, with a VRAM capacity of 192GB, lower than the market's previous expectations for a more aggressive configuration of the Rubin Ultra.


This is not the final product specification officially confirmed by NVIDIA. The Rubin GPU disclosed in NVIDIA's technical blog on July 21st is a dual compute die utilizing HBM4, offering up to 288GB, 22TB/s bandwidth, and 50 PFLOPS NVFP4 performance. This supports the fundamental judgment that "Rubin is the next-generation key product," but the 192GB mainstream version of Rubin Ultra should still be considered as the customer preview information disclosed by SemiAnalysis.


For AI data center customers, VRAM, HBM supply, single-card power, and rack deployment difficulties directly affect how training and inference systems are designed, how quickly they are built, and how costly they are. If Rubin Ultra compresses mainstream VRAM to 192GB, it appears more like NVIDIA prioritizing a version that can be delivered on a large scale more easily under limited HBM and power conditions, rather than pushing the single GPU parameters to the maximum.


192GB VRAM Lower Than Early Expectations, Mainstream Version Emphasizes Deployability First


Previously, at GTC 2025 and in various brokerage and media materials, Rubin Ultra was often described as a next-generation product with higher VRAM, higher power consumption, and a more complex package, including more aggressive concepts such as HBM4E, 1TB, NVL576, and others.


The customer preview information provided by SemiAnalysis this time is significantly more conservative: the mainstream version features 8-Hi HBM4 and 192GB of VRAM. This change should be understood as a reduction compared to previous market expectations and early roadmaps, rather than NVIDIA formally announcing a specification and then retracting it.


VRAM is not an isolated parameter. Large-scale model training and long-context inference require more high-speed memory. The larger the VRAM, the more model fragments, cache, and data a single GPU can handle. However, HBM is also one of the tightest and most expensive components in current AI hardware. Using less HBM in the mainstream version may help NVIDIA allocate limited supply to more GPUs and systems.


This is also where the concept of "downspec" can be easily misunderstood. The existing information does not directly support labeling it as an AI demand slowdown. A more reasonable explanation is that in an environment where both HBM and power are constrained, mainstream products need to strike a balance between performance, cost, supply, and deployment.


Compute Narrative Shifts to NVL576, Customers Buy Integrated Systems


SemiAnalysis stated that the theoretical peak FLOPs of the Rubin Ultra mainstream version are roughly maintained, with adjustments primarily focused on memory capacity, power consumption, and system form factor, rather than simply weakening computational capabilities.


The focus of this approach is on NVL576. In simple terms, NVIDIA is still combining more GPUs through high-speed interconnects to form a larger unified computing domain, allowing customers to transition from single-card parameters to the system capabilities of entire racks or clusters. The report mentions a planned system form factor of an interconnected 8-cabinet configuration, each containing 72 GPUs, forming the NVL576 scale.


This is more practical for cloud providers and large AI labs. What limits the expansion of training clusters is not just how powerful a single GPU is, but also factors such as cabinet power consumption, cooling, power delivery, networking, HBM supply, deployment pace, and data center construction. While higher single-card memory capacity is valuable, if it results in higher power consumption, higher BOM costs, and more difficult cabinet deployment, customers may not be able to fully utilize it.


The mainstream change in Rubin Ultra is more about preserving computational performance and system scalability, while keeping memory capacity and power consumption within a range that is easier to mass-produce, supply, and deploy.


1800W Becomes the Mainstream Power Level, Power Constrains Extreme Configurations


Power consumption is another key figure in this shift in approach.


SemiAnalysis stated that the chip-level power consumption of the Rubin Ultra mainstream version is about 1800W, similar to the previous Rubin 1800W Max-Q configuration, and lower than the previously discussed 2300W version. The report also mentioned that NVIDIA may still offer 2600W to 2800W Max-P high-power versions, and there may be a 1200W SKU aimed at lower computational workloads, such as token decoding.


These details have not yet been officially confirmed by NVIDIA and should still be viewed in the context of the report. However, they all point to the same practical issue: even if the chip itself can run at higher power levels, customers' data centers may not be able to reliably accommodate this.


The higher the power consumption of a single GPU, the higher the requirements for cabinet power supply and cooling. Many customers are not facing the question of whether they want to buy a more powerful GPU, but rather whether their existing facilities, power access, and cooling systems can support it. SemiAnalysis has previously mentioned that Rubin has two power configurations: 2300W Max-P and 1800W Max-Q, with Max-Q emphasizing performance per watt, leading customers to choose lower power operation due to power constraints.


In a power-constrained environment, the 1800W version may actually be more suitable for mass deployment. It sacrifices some of the extreme configurations but helps reduce power delivery and thermal pressures, bringing it closer to a form factor that customers can deploy more rapidly.


HBM Remains the Bottleneck, Spec Downgrade Not a Peak in Demand


The tight supply of HBM is another key driver for the memory downgrade to 192GB.


According to SemiAnalysis, HBM is relatively tighter compared to TSMC's front-end manufacturing capacity. By reducing the HBM usage per single GPU, NVIDIA can allocate the limited supply to more products and hedge against the cost pressure from future HBM price increases. This assessment also needs to be viewed in the context of the current tightness in the AI hardware supply chain.


This has a direct impact on the industry chain. If the mainline GPU's single-card HBM capacity is lower than previously expected in the market, the HBM supplier's unit demand estimation, cloud vendor procurement models, and server BOM will all undergo changes. However, this does not mean that HBM is not tight. On the contrary, it is precisely because of the tightness in HBM supply that NVIDIA is motivated to adjust the mainline solution to a more HBM-efficient and easier to mass-produce form.


The Rubin Ultra may not necessarily be left with only one version. According to SemiAnalysis, the high-power Max-P version may still be targeted at specific customers and extreme scenarios, but it may not become the most widely deployed mainline SKU.


This disclosure feels more like a product of AI compute expansion entering the engineering constraint phase. NVIDIA has not halted its push towards larger system scales, but whether the next-generation GPU can be smoothly scaled up increasingly depends on HBM supply, data center power, and customer deployment capabilities, rather than just single-chip parameters.



Recommended

Morgan Stanley Analysis: Enterprise SSD Capacity Doubles, Storage Supply Faces Test

Aug 3, 16:40
Morgan Stanley Analysis: Enterprise SSD Capacity Doubles, Storage Supply Faces Test

Kioxia's profit margin approaches 80%, JPMorgan Chase raises its target price to ¥155,000

Aug 3, 16:37
Kioxia's profit margin approaches 80%, JPMorgan Chase raises its target price to ¥155,000

Morgan Stanley 120-Page Research Report Deep Dive: How to View the AI Market Correction?

Aug 3, 16:30
Morgan Stanley 120-Page Research Report Deep Dive: How to View the AI Market Correction?

After Binance Alpha listed the US stock meme, how far can this narrative really go?

Aug 3, 16:12
After Binance Alpha listed the US stock meme, how far can this narrative really go?

Consecutive Two-Week TACO Weekend | TradeXYZ Weekend Market Watch

Aug 3, 15:53
Consecutive Two-Week TACO Weekend | TradeXYZ Weekend Market Watch

Rare U.S.-Japan Cooperation in Nearly 30 Years Marks the End of the Yen Carry Trade Era

Aug 3, 15:31
Rare U.S.-Japan Cooperation in Nearly 30 Years Marks the End of the Yen Carry Trade Era