Three Chips, One Agent: How Nvidia Split the Work of Reasoning, Talking, and Managing
Nvidia's new stack treats an AI agent's workflow as three separate jobs instead of one. The Vera Rubin NVL72 rack, built around the Rubin GPU, handles the heavy reasoning and retrieval steps agentic tasks require, while a separate Vera CPU rack manages orchestration - tracking tool calls and coordinating the thousand-step journeys a single prompt can now trigger. The Groq 3 LPX rack, inherited from Nvidia's Groq acquisition, is dedicated purely to generation: turning that reasoning into a fast, readable response. Nvidia says the full seven-chip Vera Rubin platform is now ramping into full production, with shipments to customers beginning this fall [1].
That specialization shows up directly in the benchmarks. On Gemma 4 31B with a 100,000-token context, Groq 3 LPX produced 3,400 output tokens per second - about four times faster than the nearest alternative platform for the same long-context workload [2]. The number matters less as a raw speed record than as a signal of what Nvidia thinks agentic AI actually needs: not one chip doing everything, but a pipeline where the slowest, most user-facing step gets its own purpose-built hardware.
The strategy extends past Nvidia's own silicon. NVLink Fusion opens the same rack infrastructure - NVLink components, MGX rack designs, manufacturing partners, and cluster software - to outside chipmakers building custom XPUs and CPUs, with Intel contributing x86 CPUs and Samsung Foundry handling design-to-manufacturing for the resulting third-party silicon [3]. Read together, the message is that Nvidia wants to own the agentic AI factory's plumbing even for hardware it didn't build.



