memujo
Technology7 min read

Google's Bug Bounty Math Broke. AI Slop Won.

Google paused open source bug bounty submissions on October 1 after invalid AI reports overwhelmed its engineers, and curl's data shows exactly why.

By Alice

In this article
  1. 01What Actually Changed
  2. 02The Triage Math Nobody Budgeted For
  3. 03Why This Matters Beyond Google
  4. 04Our Read

On October 1, Google stopped accepting product vulnerability reports in its Open Source Software Vulnerability Reward Program, the bounty program that pays researchers for finding flaws in Golang, Angular, Bazel, Protocol Buffers, Fuchsia, and critical third-party dependencies. The reason, in Google's own words: "This pause is due to a significant rise in automated submissions, the vast majority of which are not valid." An update on the program's future is promised for Q1 2027, per the notice on the Google Bug Hunters rules page and as TechCrunch reported.

Read as a news item, this is a story about AI slop annoying a big company. Read as an engineer who has run a review queue, it is something sharper: the first time the world's largest payer for security research conceded that its market has broken, because the cost of generating a plausible bug report fell to nearly zero while the cost of verifying one did not move at all.

What Actually Changed

The mechanics are narrow and worth stating precisely, because several outlets overstated them. The pause covers new product vulnerability submissions to the OSS VRP. It does not cover supply chain reports (flaws in repository settings, GitHub Actions, access control rules), it does not affect reports submitted before October 1, and Google says some reports against Google Cloud repositories can still route through the Cloud VRP. Bleeping Computer adds that researchers can still earn up to $15,000 per fix through the Google Patch Rewards Program, which pays for working patches rather than claims.

The volume problem is real and recent. Google launched the OSS VRP in August 2022 with payouts from $100 to $31,337 per finding. Across all its vulnerability reward programs since 2010 it has paid out more than $81.6 million, and 2025 was the record year: $17.1 million to over 700 researchers, up 40 percent from $12 million in 2024. That money was, until recently, buying genuine signal.

According to Tom's Hardware, Google engineers and open source maintainers were drowning in thousands of reports that were invalid or outright hallucinations, spending more time manually validating code than fixing real vulnerabilities. Reporting used to be painstaking manual work. LLMs and automated bug-hunting scripts collapsed the cost and effort of producing a report to near nothing.

Google is not an isolated case; it is the highest-profile exit from a trend with a growing body of receipts:

Program Action Date Numbers on record
curl (HackerOne bounty) Program ended entirely Ended Jan 31, 2026 87 confirmed vulns, $100k+ paid over 6 years; confirmed rate fell from 15%+ to under 5%
Linux kernel security list Submission limits (two/month per submitter) May 2026 Kernel nearing 2,000 CVEs per release across 40M lines; maintainers "completely overwhelmed"
Intel bug bounty (Intigriti) All financial rewards removed Mid-September 2026 Had paid up to $100,000 per flaw; no official explanation
Google OSS VRP Product vulnerability submissions paused Oct 1, 2026 $81.6M paid across Google VRPs since 2010; $17.1M in 2025 alone

The curl case is the most instructive because the maintainer published the actual funnel data. In the announcement ending curl's bounty, Daniel Stenberg wrote that the confirmed rate, historically north of 15 percent of submissions, plummeted below 5 percent starting in 2025: "Not even one in twenty was real." He also flagged something Google's notice hints at but does not spell out: even reports that were not obviously slop were lower quality, "presumably because they too were actually misled by AI but with that fact just hidden better."

The Triage Math Nobody Budgeted For

Here is the data scientist's read. A bug bounty intake queue is a classification problem with an extreme class imbalance. The base rate of true, novel, in-scope vulnerabilities in a mature codebase is low. For years the submitter-side filter was the most expensive part of the pipeline: reading the code, understanding the protocol, constructing a reachable exploit path. That labor acted as a free precision filter. When language models absorbed that labor, precision collapsed while the verification workload per report stayed constant, because verifying a hallucinated race condition in a state machine still requires a human to read the state machine.

Run curl's published numbers through that lens. Call the average human triage time per report T. At a 15 percent confirmed rate, the review labor per real vulnerability was roughly 6.7T. At a 4.9 percent confirmed rate (the midpoint of "below 5 percent"), it is about 20T. Triage cost per real find tripled with no change to the code, the researchers' skill, or the payout schedule. The only variable that moved was submission volume from a generator whose marginal cost per report is a few cents of tokens.

Now scale that to Google's program and the asymmetry becomes absurd. A single $31,337 top-tier bounty is a budgeted line item. Engineer hours spent debunking thousands of hallucinations are not. If an automated submission arrives for pennies and takes an engineer thirty minutes (a conservative, labeled assumption here, not a Google-reported figure) to disprove, the attacker-equivalent cost ratio between generating noise and clearing it is on the order of three to four orders of magnitude. No reward schedule survives that ratio, because bounties pay only on confirmed findings while triage costs accrue on every submission, real or not.

The queueing consequences follow directly. Triage capacity is fixed at headcount times focus. Submission arrival rate went from bounded-by-human-effort to effectively unbounded. A queue where arrivals grow and service does not has exactly three fates: drop submissions, degrade service, or shut the front door. Google chose the third for product reports while deliberately keeping the supply chain channel open. That carve-out is the tell. Supply chain reports (a poisoned Actions workflow, a wrong access control rule) are lower in volume, cheap to verify against repo configuration, and hard to hallucinate into existing. Google paused the channel where verification is expensive per report and kept the one where it is cheap. That is not a retreat, it is triage of the triage system.

There is a second-order problem Google cannot automate away: you cannot reliably deploy a model to filter the slop flood when the flood is generated by the same class of models. A detector tuned to catch hallucinated vulnerability reports will also reject the genuine-but-unusual report from the researcher who found the real one, and in security intake, a false negative is a disclosed zero-day in the wild. The precision-recall trade lives in exactly the place programs cannot afford to tune.

Why This Matters Beyond Google

Every security team should notice that the intake layer, not the patch layer, is what failed. Nothing about Google's ability to fix bugs degraded; the report pipeline degraded. The same pattern 100 tech companies warned about in their open letter on AI cyber threats was framed as attackers getting new capabilities, but this story is the defensive mirror image: offense and defense both got LLMs, and the side that merely has to generate plausible text has the cheaper job. The Hugging Face incident involving rogue OpenAI agent bots showed agents attacking infrastructure at machine speed; the OSS VRP pause shows what machine speed does to a queue staffed by humans on the other side.

Microsoft, for its part, warned in May that AI tools would increase the "pace and breadth of vulnerability discovery" across the industry and "can raise operational demands." Last month it shipped patches for a record 966 flaws, including two actively exploited zero-days. More real bugs found faster is a genuine win, and it is the part everyone quotes. The part nobody budgeted is that the same tools ship you 19 plausible fakes for every real finding, and a fake CVE report still consumes a human.

Expect the market to reprice what a report has to contain. The Patch Rewards model (pay for a working patch, not a claim) and the supply chain carve-out (reports verifiable against configuration, not interpretation) are two versions of the same principle: shift the burden of proof to the submitter, because verification cost is the scarce good now. Expect mandatory proof-of-concept exploits, executable reproducers, and structured evidence requirements to migrate from elite private programs into public bounty intake everywhere. The era of "please describe the vulnerability" is over.

Our Read

Our prediction, stated so it can be wrong: when Google publishes its Q1 2027 update, the reopened OSS VRP will not simply resume. It will come with evidence-first intake, meaning a runnable proof of concept or patch attached to every product vulnerability submission, and the effective bounty mix will tilt toward supply chain and patch rewards, where Google just demonstrated verification economics still work. If instead the program reopens with the same intake format as 2025, we will owe the program a rewrite, because that would mean Google found a way to staff its way back under the slop line, and nothing in the curl or Intel data suggests staffing is a durable answer to a generator whose cost trends toward zero.

The uncomfortable generalization, for anyone who runs review queues of any kind, from code review to content moderation to peer review: the reports your team receives were implicitly filtered by the effort required to produce them. That filter is gone, permanently. Whatever your queue accepts, budget for a world where the cost of sending you something is a few cents, and make the absence of evidence itself the disqualifier. Google just spent its way to that conclusion at a scale no one else can afford to test twice.

  • #google
  • #bug-bounty
  • #ai
  • #security
  • #open-source
  • #triage

Sources

Share this story