The Price of Working Harder
Google frames Gemini 3.8 Flash's benchmark gains as coming from a model that 'works harder' on complex tasks, taking extra reasoning steps and calling tools iteratively rather than answering in one pass [1]. That framing has a cost: Artificial Analysis found that running its full Intelligence Index suite on 3.8 Flash cost about 40 percent more than on 3.7 Flash - $0.58 versus $0.40 per task - driven by roughly 30 percent more output tokens and more turns on agentic evaluations [2]. Some of the improvement, in other words, is bought with token spend rather than raw capability. Reddit's r/singularity community reached a similar conclusion independently within hours of launch, noting the model was burning even more tokens than its predecessor and joking that Google is basically buying performance by letting the model spend more on tokens. The same thread raised a sharper concern: the model's headline 89.4 percent score on Terminal-bench 2.1 [3]looks far less impressive next to community reports of a much lower score on a newer, harder Terminal-bench 4.0 suite - a live, crowd-sourced check on how fast a launch-day benchmark claim gets scrutinized.



