A Deep Dive into NVIDIA's Rubin Downgrade Rumors: The standard version of Vera Rubin has been delivered to dozens of customers and is set for a large-scale release in the fall, while the specifications for the Rubin Ultra have not been finalized.
August 2nd: Analyst qinbafrank posted to clarify the market's discussion on NVIDIA's Rubin HBM downgrade, stating that the actual situation is not a downgrade of the delivered standard Vera Rubin NVL72, but rather an upgraded version, Rubin Ultra, planned for the second half of 2027 with the configuration yet to be finalized. The standard Vera Rubin is progressing smoothly, with Dell being the first to deliver the initial batch of NVL72 systems to CoreWeave in early June and completing the industry's first full-scale boot test by July. By July, dozens of customers have received test racks or early shipments, including Microsoft, OpenAI, Anthropic, Google Cloud, Oracle, Nebius, and SpaceX AI, some of which are already running in customer data centers. A larger scale delivery will begin in the fall, with overall progress ahead of Blackwell, significantly reducing rack assembly time to about 5 minutes.
The original Rubin Ultra's aggressive configuration announced at GTC 2026—featuring 4 computation chips close to the reticle limit and 16 HBM4E stacks in a single package, achieving about 1TB of memory—was questioned by SemiAnalysis and other organizations at the end of June due to TSMC's CoWoS-L substrate warpage, reticle size limitations, and yield challenges. The four-chip solution is said to have been scrapped, possibly shifting to a dual-chip design with 8 HBM4E stacks, similar to the standard version, with a single-package capacity of about 384GB. Although significantly reduced from the original target, it can partially compensate for system-level performance through rack-level expansion.
At the end of July, TrendForce further pointed out that the HBM specification for Rubin Ultra is still undecided. Against the backdrop of tight supply and ongoing price increases, NVIDIA is prioritizing shipment volume and I/O speed, considering lower-spec options including HBM4E 8-layer stacks, with the key focus on balancing capacity with supply certainty. The final specifications are expected to be finalized after validation in the second half of 2026, with NVIDIA still aiming for shipment in 2027, but the overall pace has shifted from aggressive to pragmatic.