Nvidia skips the mid-cycle refresh and bets everything on 1.6 nm feynman gpus

NVIDIA has just torn up its own rulebook. For the first time ever the company is abandoning the annual 'Super' refresh, leaving the RTX 50 series without a mid-life facelift and jumping straight to a 1.6 nm monster it calls Feynman. Jensen Huang will pull the curtain back at the March 2026 CTC conference, and the room will feel the heat—literally: dual-die cards that can sip 2 000 W are already chilling the liquid loops in Santa Clara labs.

The nanometre that changes the power meter

TSMC’s A16 node is the star. It is the first process to print transistors at 1.6 nm, shaving line widths just enough to push electron interference back into the textbook and deliver a 20 % jump in power efficiency over today’s 2 nm wafers. NVIDIA is pairing that gain with an SPR back-side power-delivery mesh, a lattice of hidden rails that keeps voltage droop away from the AI cores that will chew through trillion-parameter models. The catch: every extra ampere saved inside the silicon is spent outside it, because the die is bigger, the clocks are higher and the memory stacks are screaming for juice.

Board partners have already received the thermal brief: 1 kW is the floor, 2 kW the ceiling. Sub-ambient cooling is no longer a PR stunt; it is a spec. Data-centre racks that once hosted four A100 cards will now host one Feynman module and a radiator the size of a refrigerator. NVIDIA will swallow the complexity by borrowing Intel’s EMIB-T bridge packaging, freeing itself from TSMC’s CoWoS capacity chokepoint and letting the GPU giant mix compute tiles, I/O tiles and HBM4 stacks like Lego.

Rubin becomes a stepping stone, not a destination

Rubin becomes a stepping stone, not a destination

Industry roadmaps still list Rubin for late 2025, but inside the company the chip is already a bridge. Rubin cards will ship, reviewers will run benchmarks, and then NVIDIA will pivot the spotlight to Feynman, relegating Rubin to a single-year placeholder. The message to AMD, Intel and the Chinese upstart Lisuan is blunt: catch up while we change the track again.

Gamers will see the first fruits in the RTX 60 generation, but the real target is inference at scale. Feynman is being designed as a training engine that can pivot to low-latency LPU inference without leaving the socket. One die crunches weights at night, the same die serves users at breakfast. Cloud providers love the pitch; their accountants less so when they read the utility fine print.

And yes, the RTX 5090 fire reports are still smoking. Engineers whisper that the current Blackwell power plane was never meant to feed 600 W through a 12VHPWR connector for hours on end. Feynman solves that by moving the problem downstream: if the connector melts, at least the chip underneath will stay cool—thanks to the 800 $ water block you will pay for.

Huang will stride across the CTC stage in San Jose, black leather jacket catching the strobes, and he will promise a 3× leap in raw shader throughput. He will not mention the utility bill. But the data-centre managers in the audience will already be doing the math: a 20 % efficiency gain at the transistor, a 200 % hunger gain at the plug. The future of AI is bright, hot and thirsty. Feynman is coming, and the meter is spinning.