Specialization's double edge: baked-in silicon, a shelf life to match
Frozen v2's core bet is baking Gemini's neural architecture directly into silicon, cutting the distance data travels and the calculations needed to answer a query. Google's own engineers project 6 to 10 times more tokens per unit of power than its latest TPUs [1]- a leap that would meaningfully lower the cost of serving Gemini at scale. But the approach uses a 'flexible hardwiring' method: it locks in the model's architecture, not its weights, so Gemini's parameters can still update after the chip ships [2]. That distinction matters because the same specialization that produces the efficiency gain is also the chip's biggest liability - if Google materially changes Gemini's underlying architecture, a Frozen chip could stop being useful, which is likely why the project is reportedly viewed partly as a trial run rather than a TPU-scale production line [1].


