Small Model, Big Beatdown: The Economics of Post-Training Nemotron
NVIDIA and Palantir's new sovereign AI stack is built around Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model that only activates about 3 billion parameters per pass. What makes it notable isn't size, it's the training data: NVIDIA post-trained the model on its own captured allocation decisions, rationales and outcomes, using NeMo Anonymizer, Data Designer and AutoModel inside a governed Palantir Autopilot lifecycle [1]. On NVIDIA's internal allocation-decision benchmark, the post-trained model hit 86.7 percent accuracy, 31.2 points ahead of the untuned, far larger Nemotron 3 Ultra, and 69.2 points ahead of its own base weights [1]. At Palantir's AIPCon 11 conference, NVIDIA went further, saying the smallest Nemotron model outperformed its largest by 10 to 20 times on predictive correctness, at 95 percent lower cost than frontier models, and can be retrained on a single GPU in a couple of hours [2]. The lesson is less about scale than about codifying tacit human judgment: the model isn't smarter in the abstract, it has simply absorbed years of planners' allocation calls that a bigger, generic model never saw.


