September 2026 update to What Is a Bug Worth When a Machine Found It?
The update, stated at the strength the evidence supports. AI is no longer merely cheapening defect discovery. Frontier systems have demonstrated multi-system exploitation in real environments and, in vendor testing, hardened-target chains. But the strongest capabilities are gated. Public broker maxima and top-chain bounty anchors did not move after the July edition. The September pricing object is access to capability: who can use the model, under what safeguards, against which targets, and with what operating privileges.
Casey Ellis, September 2, 2026. Point-in-time update to the July 2026 edition. Research: gated-capability synthesis. Data: september-2026-snapshot.json.
The July thesis survives only after a material revision. The wall moved. The visible market did not—yet.
| Pricing object | September evidence | Market reading |
|---|---|---|
| Defect | Generally available discovery models got cheaper; report volume remains elevated | Supply cost down |
| Primitive | Autonomous logic and workflow exploitation is now documented in real systems | Middle-tier pressure continues |
| Chain | Real multi-system chaining; hardened-target chaining credibly claimed but not independently reproduced | Top published prices unchanged |
| Access | Trusted programs, safeguard tiers and submission sanctions now do more allocation work | Binding scarcity |
On 1 September OpenAI designated Astra its first model at the company's Critical cybersecurity threshold. OpenAI reports 100 percent on ExploitBench, two previously unknown vulnerabilities used in a chain, a hardened-browser sandbox escape to host command execution and hardened-OS privilege escalation to root. vendor-run
Astra was not public on the date of this update and no system card or independent hardened-target reproduction was available. The evidence boundary moved, but the July edition's exact public-testability falsifier did not fire. OpenAI: Path to Astra
Hugging Face's forensic timeline records about 17,600 agent actions across about 6,280 activity clusters, movement from an evaluation sandbox into Hugging Face, node-root and cluster-admin access, and lateral movement across multiple clusters in under 13 hours. realised OpenAI supplied the model and infrastructure context. METR independently reviewed more than 70,000 messages/files and about 1,300 raw transcripts, corroborating large-scale autonomous coordination while explicitly not validating every exploit detail. multi
Wiz separately documented an autonomous agent finding and exploiting a GitHub Actions injection flaw in a Snowflake repository within five days, adjusting a failed payload and reaching an internal Jira credential. single It is a real primitive-to-access chain in CI/CD, not a hardened consumer chain. Wiz
The strongest September market signal is not a changed bug price. It is a changed rule about who gets the capable system.
The model is not public at the capability level doing the economic work in this update. org's own
Google allocates the model through identity and operating standards rather than a public per-vulnerability market. org's own
Fable 5.1 is generally available for vulnerability discovery but not exploit development; Mythos reserves more permissive cyber use for trusted access. Anthropic reports about 25 percent lower typical workload cost and as much as 45 percent lower highly agentic workload cost for Fable. org's own
Google prices generally available Gemini 3.8 Flash at $0.75 per million input tokens org's own and $3.75 per million output tokens org's own, while restricting the cyber-specialised variant. Published model pricing confirms supply-cost compression; Google's cyber performance numbers remain unreplicated vendor and partner claims.
Discovery and remediation supply gets cheaper. Chain-capable operation gets rationed.
Current owner pages were rechecked on 2 September. Neither checked broker added an AI or LLM category.
| Market anchor | Current published maximum | Change since July |
|---|---|---|
| Crowdfense mobile zero-click full chain | $7M org's own offer | None found |
| Crowdfense Chrome / Safari one-click full chain | $2–3M / $2.5–3.5M org's own offer | None found |
| Operation Zero mobile / virtualization | $2.5M / $1M org's own offer | None found |
| Apple zero-click network-to-kernel | $2M org's own offer | None found |
| Google Pixel Titan M2 chain with persistence | $1.5M org's own offer | None found |
| GitHub public / VIP critical | $10K / $30K+ org's own offer | No outcome data |
These are advertised maxima, not realised transactions. No credible public sale or payout was located that would show a frontier chain clearing below the 2024–26 anchors. Absence of public evidence is not evidence that a private market did not move.
Interviews collected by Dark Reading report roughly doubled HackerOne volume, a 450 percent year-over-year ZDI peak that later moderated, and a short Bugcrowd surge above 300 percent that normalized around twice historical volume. The same reporting describes pressure in the roughly $2,000–$50,000 tier multi while HackerOne says aggregate H1 payments and the count of researchers earning $100,000 each rose 25 percent. org's own
More findings can clear at lower average unit value while total program spend and top researcher earnings still rise.
Apple raised top chain ceilings while reducing macOS TCC and sandbox awards in the same 2025 restructure. A full TCC bypass moved from about $30,500 to $5,000 multi and a macOS-only sandbox escape from about $10,500 to $5,000. multi Apple attributed the top-end increase to mercenary-spyware defense, not AI. This supports object-level sorting while weakening a monocausal AI story.
Apple's current rules say repeated ineligible or unvalidated AI-assisted reports can trigger a 180-day processing pause and eventual removal. org's own The price of low-signal supply is becoming lost access, not merely a smaller check.
Contrast reports 5 percent agreement among three AI scanners and 17 percent repeatability for one scanner across runs. It estimates $315 thin in model-token cost to scan two million lines and $128,000 thin to triage the output. The accessible methods are incomplete and the vendor sells the alternative it recommends, so the figures are directional, not dispositive.
A Splunk practitioner reports product-aware triage up to 20 times faster while retaining evidence gates and human review. thin It is a self-report, not an external benchmark, but it prevents a simple "cheap discovery means permanently expensive triage" conclusion.
VulnCheck found 14 confirmed exploited vulnerabilities among 1,061 attributed to AI-assisted discovery: 1.3 percent, roughly the overall first-half rate. single Of more than 23,000 Anthropic Project Glasswing findings, it found 126 published CVEs and one confirmed exploited vulnerability. This is the strongest available counter-evidence to treating discovery volume as a direct proxy for offensive value.
| July falsifier | Status | Why |
|---|---|---|
| Broker publishes an AI/LLM category with a price | Not triggered | Neither checked broker added one |
| Broker cuts chain prices because of automation, or a realised sale clears below prior anchors | Not triggered | Public maxima unchanged; no credible realised transaction located |
| Public system reproduces withheld hardened-target results | Near miss | Astra is gated and vendor-run; Hugging Face is real but a different target class |
| Apple or Google cuts top-chain rewards | Not triggered | The top anchors remain |
| A cut program proves triage cost made payouts binding | Not triggered | Contrast is not a cut program and methods are incomplete |
| GitHub's post-restructure signal remains unchanged | Unresolved | No public outcome series |
Strong for multi-system logic and configuration chains. For hardened browser and OS chains, the support is still vendor-run and not independently reproduced.
Supported in parts of the middle tier. Contradicted by unchanged top anchors and rising aggregate bounty payments. Private realised broker prices remain opaque.
OpenAI, Google and Anthropic independently converged on trusted or restricted cyber access while generally available discovery products became cheaper.
Explains access sanctions, submission limits and some middle-tier compression. It is incomplete because triage itself is automating and the strongest systems remain deliberately scarce.
The machine discount now applies to more than defects and primitives, but not uniformly. Multi-system chaining is real, hardened-target chaining is credibly claimed, and both remain separated from the public market by access controls, evaluation opacity and target class.
This point-in-time pass audited 27 saved sources. The machine-readable snapshot carries the evidence class and caveat for every load-bearing row.
Full synthesis: research/2026-09-02-gated-capability/RESEARCH.md. Structured observations: data/september-2026-snapshot.json.
This is a 2 September snapshot, not a September retrospective. Public maxima are not realised prices. Vendor benchmark claims remain labeled vendor claims. Absence of a public transaction is not proof that a private market did not move.