The Hallucination Problem: Why AI Slop Fools Human Reviewers
Google's own framing undercuts the easy narrative that AI-written reports are the problem by definition. An analysis of the pause argues the issue isn't that generative AI wrote the submissions - it's that so many of them are wrong. Poorly written, incomplete, and hallucinated reports are not always immediately recognizable as such [1], which is exactly what makes them expensive: a fabricated exploit path or a claim about a function that isn't reachable the way the report describes still has to be manually traced through source code before a reviewer can rule it out. Google confirmed the trigger was blunt volume, not just bad luck - a significant rise in automated submissions, the vast majority of which are not valid [2]. The two problems compound: more submissions mean more hours spent disproving plausible-sounding fictions, and that reviewer time comes directly out of time that would otherwise go toward real vulnerabilities in projects like Go, Angular, Bazel, and Fuchsia. Google's own public framing matched that math rather than panic - the announcement read as a routine administrative notice rather than a crisis statement - but outside commentary and press pickup read it more dramatically, framing it as evidence that AI slop is now breaking bug bounty programs outright, and drawing the most debate of any single post on the story.


