A trillion parameters on a fraction of the compute
Mistral trained ML4 from scratch on roughly 3,800 to 4,000 Nvidia Grace Blackwell GPUs in its own European data centers over about two months[1]. VP of Science Pierre Stock frames that as a deliberate efficiency story, saying the run used 'two to three times less' compute than Chinese competitors and 'significantly less' than closed-source Western labs[2]. That framing matters because it is Mistral's whole pitch: rather than out-spend OpenAI or Google on raw compute, the company is betting that architecture and training technique can close the gap with a fraction of the GPU budget. It also doubles as a sovereignty argument - training entirely inside Europe, on hardware Mistral controls, without leaning on a foreign hyperscaler's cloud. Whether 'less compute, comparable results' holds up once the full weights are independently benchmarked is the open question the company has essentially staked its credibility on.


