← All digests
✦ AI News for Builders

Google pauses its open-source bug bounty over AI-generated reports, Cantina's open apex-flash-1 solves 40 of 60 held-out security tasks for $2.38, Google's RRSI keeps self-improving agent harnesses from memorizing their tests, and ChatGPT ads get a visual format

Monday, October 5, 2026·8 min read·4 stories

Monday's theme is signal versus volume. AI made it cheap to file a bug report, so Google stopped paying for them. The same week, an open model shows what verified AI security work looks like, Google publishes a method for agents that tune themselves without cheating on the test, and OpenAI adds a new place for ads to live.

Story I

Google stopped taking open-source bug bounty reports because most of the AI-generated ones were wrong

Google's Open Source Software Vulnerability Reward Program stopped accepting new product vulnerability submissions on October 1. The company's reason, quoted by TechCrunch and Infosecurity Magazine: "a significant rise in automated submissions, the vast majority of which are not valid." Google says it will reformat the program and give an update in the first quarter of 2027. Supply-chain reports and anything filed before October 1 are not affected, and researchers are pointed to Google's other VRPs and the Patch Rewards Program in the meantime.

This is not a small program. OSS VRP launched in August 2022 and paid from $100 to $31,337 depending on severity and project. Neither article gives a count of how many reports came in, so we can't tell you how bad the flood got. What we can say is that a bounty program built to reward finding bugs has decided the cost of reading AI-written reports is higher than the value of the real ones mixed in.

For builders

If you run an AI scanner over open-source code, don't file anything you haven't reproduced: attach a working proof of concept and the exact commit, or don't send it. If you maintain a project, add a line to SECURITY.md today that reports without a reproducible PoC will be closed unread. It sounds harsh. Google just showed that the alternative is closing the door for everyone.

Story II

Cantina's open apex-flash-1 solves 40 of 60 held-out security tasks for $2.38. Opus 5 solves 43 for $74.68

Cantina Security released apex-flash-1, a 321B model fine-tuned from GLM-5.3-Flash under the MIT license. On its held-out evaluation it solved 40 of 60 tasks, a 66.7% pass@1, for $2.38 across the whole run. The base GLM-5.3-Flash scored 60.0% for $4.56, and Claude Opus 5 scored 71.7% for $74.68. That's roughly 31 times cheaper for five fewer points, which is the trade the model card is selling.

The card is clear about the job: a "focused worker under a larger agent's direction" for code analysis, tool use, exploit development and verification in isolated environments, served through vLLM or SGLang with OpenAI-compatible APIs. The evaluation covers text-only security tasks. There's also an "abliterated" derivative with modified refusal behavior that Cantina labels experimental. And 60 tasks is a small benchmark built by the people who trained the model, so treat the percentages as a first data point, not a ranking.

For builders

If you already pay for a frontier model to triage security findings, run a split test this week: let apex-flash-1 do the first pass on a batch of known-fixed bugs from your own repo, and send only its confirmed hits to the expensive model. Keep it in a sandbox with no network access to production, as Cantina intends. And read the previous story before you file anything it finds.

Story III

Google Research's RRSI keeps self-improving agent harnesses from memorizing their benchmarks, and the code is Apache 2.0

When an agent rewrites its own harness (prompts, tools, control flow, context handling) against a fixed set of tasks, it tends to learn the tasks instead of the skill. Google's RRSI, short for Regularized Recursive Self-Improvement, attacks that with two constraints: an edit budget that shrinks over rounds so late changes are small and traceable, and a strict critic that rejects hardcoded task-specific fixes and strips costly components that don't earn their keep.

The README reports Terminal-Bench 2.1 going from 74.2% to 80.2%, EngDesign from 50.0% to 54.9%, and a smaller gain on a workspace task. The Decoder adds that it was the only method tested that consistently stayed above baseline on five unseen benchmarks, with gains of up to 4.7 points there, while using about 30% fewer tokens at runtime. The main experiments kept Claude Opus 4.8 frozen, and a harness tuned with Gemini 3.5 Flash transferred to Gemini 3.1 Flash Lite unchanged.

For builders

You don't need the repo to steal the idea. If you tune prompts or agent configs against an eval set, hold out a third of the tasks and never look at them until the end. Cap each tuning round to one or two changes so you can tell which one helped. And grep your prompts for instructions that only make sense for a specific test case. That's the hardcoding RRSI's critic throws out.

Story IV

ChatGPT ads get a visual format during image generation, plus conversion APIs through AppsFlyer, Adjust and others

OpenAI announced a visual ad format that shows labeled product images while ChatGPT is generating an image for the user, kept separate from the image itself. Testing starts later this month in the US with a first group of advertisers. Alongside it, OpenAI added conversion data integrations with Hightouch, Tealium and LiveRamp, and attribution support with AppsFlyer, Triple Whale, Adjust, Branch, Singular, Kochava and others. Self-serve signup is at ads.openai.com.

The early numbers come from partners, not OpenAI: DV Rockerbox says WeightWatchers' attributed cost per acquisition was 15.3% lower than its blended paid-search benchmark, and Triple Whale says 93% of Portland Leather's ChatGPT visitors were new. Those are vendor case studies, so discount them accordingly. Per OpenAI's help center, ads show only on Free and Go plans, never on Plus, Pro or business tiers, and advertisers get aggregated performance data, not chats.

For builders

If you sell a product and already use one of those attribution tools, the integration work is mostly done; check whether your MMP account lists ChatGPT Ads as a source before you spend a dollar. If you build on ChatGPT for your own work, open Settings, then Ad Controls, and turn off ad personalization so past chats stop feeding ad selection. Ads still match the current thread either way.

Sources

  1. TechCrunch — Google froze its open source bug bounty program (Oct 1 pause, Google's statement)
  2. Infosecurity Magazine — what's paused vs still accepted, Q1 2027 update, $100–$31,337 range
  3. MarkTechPost — apex-flash-1 release and 31x cost comparison
  4. Hugging Face — apex-flash-1 model card (base model, license, 40/60, costs, serving)
  5. The Decoder — RRSI unseen-benchmark gains, token savings, models used
  6. GitHub — google-research/rrsi README (method, results, Apache 2.0)
  7. OpenAI — new visual ad format and measurement partners
  8. OpenAI Help Center — which plans see ads, Ad Controls, advertiser data access

— The Vibe Gate news desk. We read the firehose so you can keep building.

← All digests  ·  The blog