Nvidia spent most of 2025 selling every Blackwell GPU it could build. In July 2026 the story has shifted: the company is now telling customers that its next-generation Vera Rubin platform is moving from samples into real production, and the message underneath is bigger than a single chip. The AI hardware market is pivoting from training ever-larger models to running them cheaply, at scale, forever — and Rubin is the first Nvidia system designed for that world from the silicon up.
In mid-July 2026, Nvidia moved to reassure investors and hyperscale buyers that Rubin remains on schedule for volume shipments in the second half of the year, with CEO Jensen Huang stating the platform is “already in production.” That reassurance matters because the entire AI infrastructure trade is priced on it. Here is what the platform actually delivers, and why the “inference flip” makes it the most consequential launch of the year.

What Nvidia Actually Unveiled
Vera Rubin is not one product but a rack-scale system named after the astronomer whose work provided early evidence for dark matter. The headline configuration is the Vera Rubin NVL72: a single liquid-cooled rack that fuses 72 Rubin GPUs with 36 Vera CPUs, stitched together by Nvidia’s sixth-generation NVLink fabric. Jensen Huang has framed it as “a generational leap,” and the numbers behind that phrase are unusually concrete this time.
Each Rubin GPU delivers roughly 50 petaflops of NVFP4 inference compute, so a full NVL72 rack reaches about 3.6 exaflops for inference and 2.5 exaflops for training. Nvidia positions that as up to five times the inference performance and roughly 3.5 times the training performance of the Blackwell GB200 systems shipping today. For an industry that measures progress in generational doublings, a 5x jump in a single step is aggressive.
Inside the Rubin GPU and Vera CPU
The specification that AI buyers will care about most is memory. Every Rubin GPU carries 288GB of HBM4 with 22 terabytes per second of bandwidth, adding up to more than 20 terabytes of high-bandwidth memory across a single rack. Memory bandwidth — not raw transistor count — has become the true ceiling on how fast large models can run, and HBM4 is the component that lifts it.
The Vera CPU is Nvidia’s own design rather than an off-the-shelf part: 88 custom Arm-based cores capable of tracking 176 threads, linked to each Rubin GPU by an 1.8 terabyte-per-second NVLink chip-to-chip connection. That tight CPU-to-GPU coupling is the point. Agentic and reasoning workloads bounce constantly between general-purpose logic and accelerated math, and the older model of a commodity CPU bolted to a GPU over a slow bus has become a bottleneck. Rubin treats the CPU and GPU as one coherent engine.

Built for Inference, Not Just Training
The most important thing about Rubin is not any single number — it is the workload it was built for. Through 2023 and 2024, spending was dominated by training: the one-time, capital-intensive job of teaching a model. In 2026 the balance has tipped. Every chatbot reply, every coding assistant, every autonomous agent taking an action is an inference call, and those calls now run billions of times a day. Analysts describe this crossover as the “inference flip,” the point at which the money spent running models overtakes the money spent building them.
Rubin is engineered for that economics. Nvidia claims the platform can cut the cost per token for mixture-of-experts inference by as much as ten times versus Blackwell, and train the same class of models with a quarter of the GPUs. When a company’s AI bill is dominated by inference, a 10x reduction in cost per token is not a spec-sheet flourish — it is the difference between a product that loses money on every query and one that scales profitably.
The Competitive Squeeze: AMD, ASICs, and Cost per Token
Nvidia is not launching Rubin into an empty field. AMD has detailed its own rack-scale answer built on the Instinct MI355X, promising multi-exaflop FP4 performance and positioning aggressively on memory capacity and price. More quietly threatening are the custom ASICs from Broadcom and Marvell that power Google’s TPUs, Amazon’s Trainium, and Microsoft’s Maia — silicon designed for one company’s workloads and growing far faster than the general-purpose GPU market. Industry forecasts put custom-ASIC growth near 45% a year against roughly 16% for GPUs.
Rubin is Nvidia’s argument that a general-purpose platform can still win on total cost of ownership. By driving cost per token down an order of magnitude and bundling networking, CPUs, and software into one rack, Nvidia is betting that most buyers would rather buy a finished supercomputer than assemble their own from custom parts. The pressure from ASICs is precisely why that cost-per-token claim sits at the center of the pitch.

Why It Matters
For anyone running or funding AI, Rubin sets the reference point for the next 18 months of budgets. Cloud pricing, model margins, and the return-on-investment math that underpins hundreds of billions of dollars in data-center spending all key off how much it costs to serve a token — and Rubin is designed to reset that number. If Nvidia delivers the claimed efficiency at volume, it hardens the company’s lead just as rivals were closing in. If the ramp slips or the real-world gains fall short of the marketing, an entire market priced for perfection has room to fall. The launch is as much a financial event as a technical one.
The Bottom Line
Vera Rubin is the clearest signal yet that AI hardware has entered its inference era. The differentiators are no longer just faster training runs but memory bandwidth, CPU-GPU coherence, and above all the cost of running models at planetary scale. Watch three things over the coming months: whether volume shipments actually land in the second half of 2026, whether independent benchmarks confirm the 5x and 10x claims, and how aggressively AMD and the custom-ASIC camp respond on price. The chip is impressive on paper; the economics it reshapes are what will matter.
Questions for Our Readers
We would like to hear from you: Is your organization already budgeting for the shift from training costs to inference costs, and how is it changing your AI plans?
Would a 10x reduction in cost per token change which AI features you consider viable to ship?
And do you see custom ASICs from the hyperscalers eventually displacing general-purpose GPUs in your own stack — or is a finished, integrated platform like Vera Rubin still the safer bet?
Share your thoughts in the comments.
Reference Sites
- NVIDIA Newsroom — NVIDIA Vera Rubin Platform
- Tom’s Hardware — Nvidia Launches Vera Rubin NVL72 at CES
- ServeTheHome — NVIDIA Launches Next-Generation Rubin AI Compute Platform
- TechRadar — AMD MI355X Rack vs. Nvidia Vera Rubin
- Yahoo Finance — Nvidia’s Rubin Reassurance Protects a Bigger AI Bet
Researched and written by: Peter Jonathan Wilcheck and Ray Anderson
Post Disclaimer
The information provided in our posts or blogs are for educational and informative purposes only. We do not guarantee the accuracy, completeness or suitability of the information. We do not provide financial or investment advice. Readers should always seek professional advice before making any financial or investment decisions based on the information provided in our content. We will not be held responsible for any losses, damages or consequences that may arise from relying on the information provided in our content.



