Skip to content

Vera Rubin, and why a chip you'll never buy sets next year's AI bill

Nvidia's Vera Rubin platform promises ten times the inference throughput per watt of Blackwell. No NZ enterprise will buy a rack, but the rack sets the price of every AI call you make next year. How to read a chip launch as a budget signal.

At GTC in March, Nvidia unveiled Vera Rubin: seven new chips, five rack types, and a claim of up to ten times the inference throughput per watt of Blackwell at a tenth of the cost per token. You will never buy one. A Vera Rubin NVL72 rack is hyperscaler equipment, and the first partner systems do not ship until the second half of the year. But that rack sets the price of every AI call your business makes in 2027, which makes it a line in your budget whether you noticed the launch or not. Here is how to read it, and what to do with a price that keeps falling.

What you need to know

  • The generation is real and it is in production. Jensen Huang said at CES in January that the Rubin platform was in full production, and at GTC on 16 March Nvidia said all seven chips were. Partner products arrive in the second half of 2026.
  • The number that matters is per watt, not per chip. Nvidia claims up to 10x higher inference throughput per watt and one-tenth the cost per token compared with Blackwell. Power is the scarce input in an AI data centre, so tokens per watt is what sets the price of a call.
  • The software half of the jump shipped the same day. Dynamo 1.0, Nvidia's open-source inference layer, went to general availability at GTC with a claimed 7x boost on Blackwell, and AWS, Azure, Google Cloud and Oracle have all adopted it.
  • "Wait until it's cheaper" never arrives. Nvidia is guiding to at least US$1 trillion in revenue from 2025 through 2027. Each generation lands roughly yearly and each one drops the cost of a token again. There is no bottom to wait for.
  • Budget for the workflow and the data, not the model. The model and the inference behind it are the parts getting cheaper on someone else's capital. What you pay for once and keep is the work you have wired AI into.

10x

Inference throughput per watt claimed for Vera Rubin over Blackwell, at one-tenth the cost per token

Source: Nvidia, March 2026

288GB

HBM4 memory on a Rubin GPU, with 50 petaflops of NVFP4 inference compute, 5x Blackwell

Source: ServeTheHome, January 2026

US$1T

Minimum revenue Nvidia says it expects from 2025 through 2027, up from US$500 billion guided a year earlier

Source: Nvidia; eWeek, March 2026

What was actually announced

Two announcements, three months apart. At CES on 5 January, Huang introduced Rubin as a six-chip platform and said it was in full production. The GPU carries 288GB of HBM4 memory and is rated at 50 petaflops of NVFP4 inference and 35 petaflops of training, which Nvidia framed as 5x and 3.5x Blackwell. Read that carefully, because "full production" in January meant the chips were back from the fab and systems were being brought up in Nvidia's labs. Nothing was shipping.

At GTC on 16 March the platform grew to seven chips and five rack types: the Vera Rubin NVL72 GPU rack with 72 Rubin GPUs and 36 Vera CPUs, a Vera CPU rack, a Groq 3 LPX inference accelerator rack, BlueField-4 storage and Spectrum-6 Ethernet. The headline claims are the ones a CFO should keep: up to 10x higher inference throughput per watt at one-tenth the cost per token compared with Blackwell, and training a large mixture-of-experts model with a quarter of the GPUs. Groq 3 LPX is rated at up to 35x higher inference throughput per megawatt. Partner systems ship in the second half of the year.

The same day, Dynamo 1.0 went to general availability. It is the software that splits an inference job across a cluster so the hardware is never idle, and Nvidia says it lifts Blackwell inference by up to 7x. Every major cloud has adopted it, which matters more for your bill than the chip does, because it lands on hardware that already exists.

Why a rack you will never buy is your problem

Nobody in New Zealand is putting a Vera Rubin NVL72 in a comms room. The rack is liquid cooled and bought by hyperscalers in gigawatt lots. That is exactly why it matters to you. Your AI runs on a model, the model runs on a cloud, and the cloud runs on these racks. When the cost of a token drops by an order of magnitude at the bottom of that stack, the price you pay at the top follows, with a lag of a year or so.

This is the mechanism that has held for three generations. Hopper made large models practical, Blackwell made reasoning models affordable, and Rubin is being sold on tokens per watt because power, not silicon, is now the constraint on the supply side. When a vendor competes on efficiency instead of raw speed, the customer sees it as a falling price per call. The frontier model that costs you a certain amount per million tokens today will cost a fraction of that on the same hardware budget once Rubin is in the clouds, and the cheaper tier below it will do what the frontier does now.

So a chip launch is a budget signal. Read it as one. The specific figures will move, and Nvidia's claims are Nvidia's claims until an independent benchmark confirms them. But the direction is not in doubt, and the direction is what a budget should be built on.

What to budget for instead

Here is the part that changes a plan. If inference is on a falling schedule that you do not control and did not pay for, then the model is the wrong thing to make your AI budget about. It is a variable cost that trends down. The assets that hold their value are the ones the chip does nothing for: your data, clean and connected; the workflows you have wired AI into; and the governance that lets you swap models without a risk review each time.

That is why we built RIVER AI to be model-agnostic. When a Rubin-class price drop lands, a client on our command centre gets it by changing a setting, not by scheduling a project. The money went into the workflows and the data. Those compound. The model underneath is a swappable part that keeps getting cheaper.

The trap is the opposite plan: an AI budget that is mostly model spend, tuned to one vendor's pricing page, with nothing owned underneath. When the price falls that plan does not get cheaper. It gets rebuilt.

I read a chip launch the way I read a fuel price forecast. I am never going to buy a refinery, but the forecast tells me what to budget for the fleet. Rubin says the cost of running a workflow on a frontier model keeps dropping on a schedule I do not control. So I put the money into the workflows and the data, and I let Nvidia's capital pay for the rest.

Isaac RolfeManaging Director

What to actually do

Model your 2027 AI spend with a falling inference price, not a flat one. Take what a workflow costs to run today and plan on the same workload costing a fraction of that once Rubin systems are in the clouds. If a case only works at next year's price, that is a reason to build now on a smaller model, not a reason to wait.

Decide now where the saving goes. A falling price per call either shrinks your AI line or buys more workflows for the same money. The second is the growth plan: put the freed budget into the next process to wire up, and into the data that all of them depend on. Keeping the model in config and owning your evals is what makes the swap cheap, and A new model every fortnight covers that.

Read the next launch the same way. Chips, racks, and inference software will keep landing on a yearly cadence. Skip the teraflops, find the per-watt and cost-per-token figure, and ask what it does to the price of a call. That is the whole signal.