Mar 16, 2026 · View original article
Nvidia puts the Vera Rubin platform into full production at GTC 2026
At GTC on 16 March 2026 Nvidia announced that all seven chips of its Vera Rubin platform are in full production, promising large gains in inference per watt and shipping through cloud and OEM partners in the second half of 2026.
At its GTC conference in San Jose on 16 March 2026, Nvidia announced that the Vera Rubin platform, the successor to Blackwell, is in full production. The company describes it as a rack-scale AI supercomputer built from seven chips: the Vera CPU, the Rubin GPU, the NVLink 6 switch, the ConnectX-9 SuperNIC, the BlueField-4 DPU, the Spectrum-6 Ethernet switch and, newly, the Groq 3 LPU, reflecting Nvidia's integration of Groq's inference technology.
Nvidia's performance claims focus on agentic inference rather than raw training. It says a Vera Rubin NVL72 rack can train large mixture-of-experts models with a quarter of the GPUs needed on Blackwell, and deliver up to ten times the inference throughput per watt at a tenth of the cost per token. The Groq 3 LPX component is credited with up to 35 times more inference throughput per megawatt for latency-sensitive workloads, BlueField-4 storage with up to five times higher inference throughput, and a power-management scheme called DSX Max-Q with fitting 30% more infrastructure into a fixed power envelope.
The customer list underlines how concentrated the market has become. AWS, Google Cloud, Microsoft Azure and Oracle are named as cloud partners; Cisco, Dell, HPE and Lenovo as system builders; and Anthropic, Meta, Mistral AI and OpenAI as frontier labs that will run on the platform. Availability is slated for the second half of 2026. Nvidia did not disclose order values in the announcement, though CEO Jensen Huang has spoken publicly about visibility into hundreds of billions of dollars of Blackwell and Rubin demand.
The context is a build-out that shows no sign of cooling. Stanford's AI Index would report a month later that global AI investment exceeded $580 billion in 2025, and hyperscaler capital expenditure guidance for 2026 remains at record levels. What is changing is the workload mix: agents that run for hours, call tools thousands of times and consume long contexts shift the economics from training to inference, which is precisely where Nvidia is pitching Rubin. Power, not silicon, is increasingly the binding constraint, hence the emphasis on throughput per watt and per megawatt.
For enterprises the direct question is rarely whether to buy an NVL72 rack. It is how the next generation of cloud pricing, regional capacity and model availability will be shaped by where these systems land first.
What it means for leaders
- Model your inference costs, not just training. Long-running agents make token throughput and latency the dominant line items; ask providers when Rubin-class capacity reaches your regions and what it does to unit prices.
- Bake energy and location into AI governance. NIST AI RMF's Map function and ISO/IEC 42001 impact assessments should cover data-centre power, water and jurisdiction, which will drive both cost and regulatory exposure.
- Reassess concentration risk. With four clouds and four labs on one hardware roadmap, a supply hiccup ripples widely; keep multi-provider fallbacks for critical AI workloads.
- Plan capacity for on-prem regulated workloads. If data-residency rules push you to private deployments, order lead times for OEM Rubin systems will matter more than list price.
- Treat vendor performance claims as claims. Validate throughput and cost figures against your own workloads before committing to multi-year contracts.
