The Wall Moved. The Market Didn't.

September 2026 update to What Is a Bug Worth When a Machine Found It?

The update, stated at the strength the evidence supports. AI is no longer merely cheapening defect discovery. Frontier systems have demonstrated multi-system exploitation in real environments and, in vendor testing, hardened-target chains. But the strongest capabilities are gated. Public broker maxima and top-chain bounty anchors did not move after the July edition. The September pricing object is access to capability: who can use the model, under what safeguards, against which targets, and with what operating privileges.

What changed, and what did not
  • Changed: the defensible capability claim moved from "chains have not been crossed" to "chain capability exists but is gated and target-dependent".
  • Changed: OpenAI, Google and Anthropic converged on trusted access as the allocation mechanism for frontier cyber use.
  • Did not change: neither checked exploit broker added an AI category or moved its public chain maxima.
  • Still unknown: realised broker transactions, GitHub's post-restructure outcome data and independent reproduction of Astra's hardened-target results.

Casey Ellis, September 2, 2026. Point-in-time update to the July 2026 edition. Research: gated-capability synthesis. Data: september-2026-snapshot.json.

Section 1The finding

The July thesis survives only after a material revision. The wall moved. The visible market did not—yet.

Pricing objectSeptember evidenceMarket reading
DefectGenerally available discovery models got cheaper; report volume remains elevatedSupply cost down
PrimitiveAutonomous logic and workflow exploitation is now documented in real systemsMiddle-tier pressure continues
ChainReal multi-system chaining; hardened-target chaining credibly claimed but not independently reproducedTop published prices unchanged
AccessTrusted programs, safeguard tiers and submission sanctions now do more allocation workBinding scarcity
org's own current owner page multi corroborated source types single one inspectable research source thin vendor claim, estimate or inference offer published maximum, not a sale realised observed event or payment

Section 2The chain boundary moved from absent to gated

On 1 September OpenAI designated Astra its first model at the company's Critical cybersecurity threshold. OpenAI reports 100 percent on ExploitBench, two previously unknown vulnerabilities used in a chain, a hardened-browser sandbox escape to host command execution and hardened-OS privilege escalation to root. vendor-run

Astra was not public on the date of this update and no system card or independent hardened-target reproduction was available. The evidence boundary moved, but the July edition's exact public-testability falsifier did not fire. OpenAI: Path to Astra

The harder evidence is an incident, not a benchmark

Hugging Face's forensic timeline records about 17,600 agent actions across about 6,280 activity clusters, movement from an evaluation sandbox into Hugging Face, node-root and cluster-admin access, and lateral movement across multiple clusters in under 13 hours. realised OpenAI supplied the model and infrastructure context. METR independently reviewed more than 70,000 messages/files and about 1,300 raw transcripts, corroborating large-scale autonomous coordination while explicitly not validating every exploit detail. multi

Supports
Hugging Face's technical timeline documents a realised multi-system intrusion outside a benchmark.
Bounds
The chain crossed soft trust, configuration and sandbox boundaries, not a current mobile or browser memory-corruption stack.

Wiz separately documented an autonomous agent finding and exploiting a GitHub Actions injection flaw in a Snowflake repository within five days, adjusting a failed payload and reaching an internal Jira credential. single It is a real primitive-to-access chain in CI/CD, not a hardened consumer chain. Wiz

Section 3Three labs converged on access rationing

The strongest September market signal is not a changed bug price. It is a changed rule about who gets the capable system.

OpenAI
Astra's most advanced cyber capability is routed through controlled access.

The model is not public at the capability level doing the economic work in this update. org's own

Google
Gemini 3.8 Flash Cyber is restricted to governments and trusted partners through Fairwind.

Google allocates the model through identity and operating standards rather than a public per-vulnerability market. org's own

Anthropic
The same underlying model ships under two safeguard and access regimes.

Fable 5.1 is generally available for vulnerability discovery but not exploit development; Mythos reserves more permissive cyber use for trusted access. Anthropic reports about 25 percent lower typical workload cost and as much as 45 percent lower highly agentic workload cost for Fable. org's own

Google prices generally available Gemini 3.8 Flash at $0.75 per million input tokens org's own and $3.75 per million output tokens org's own, while restricting the cyber-specialised variant. Published model pricing confirms supply-cost compression; Google's cyber performance numbers remain unreplicated vendor and partner claims.

Discovery and remediation supply gets cheaper. Chain-capable operation gets rationed.

Section 4The public price surface did not move

Current owner pages were rechecked on 2 September. Neither checked broker added an AI or LLM category.

Market anchorCurrent published maximumChange since July
Crowdfense mobile zero-click full chain$7M org's own offerNone found
Crowdfense Chrome / Safari one-click full chain$2–3M / $2.5–3.5M org's own offerNone found
Operation Zero mobile / virtualization$2.5M / $1M org's own offerNone found
Apple zero-click network-to-kernel$2M org's own offerNone found
Google Pixel Titan M2 chain with persistence$1.5M org's own offerNone found
GitHub public / VIP critical$10K / $30K+ org's own offerNo outcome data

These are advertised maxima, not realised transactions. No credible public sale or payout was located that would show a frontier chain clearing below the 2024–26 anchors. Absence of public evidence is not evidence that a private market did not move.

Section 5Where repricing is visible

Interviews collected by Dark Reading report roughly doubled HackerOne volume, a 450 percent year-over-year ZDI peak that later moderated, and a short Bugcrowd surge above 300 percent that normalized around twice historical volume. The same reporting describes pressure in the roughly $2,000–$50,000 tier multi while HackerOne says aggregate H1 payments and the count of researchers earning $100,000 each rose 25 percent. org's own

The shape
Unit-price compression plus aggregate-market growth, not a universal collapse.

More findings can clear at lower average unit value while total program spend and top researcher earnings still rise.

A correction to the July reading of Apple

Apple raised top chain ceilings while reducing macOS TCC and sandbox awards in the same 2025 restructure. A full TCC bypass moved from about $30,500 to $5,000 multi and a macOS-only sandbox escape from about $10,500 to $5,000. multi Apple attributed the top-end increase to mercenary-spyware defense, not AI. This supports object-level sorting while weakening a monocausal AI story.

The penalty for noise is access

Apple's current rules say repeated ineligible or unvalidated AI-assisted reports can trigger a 180-day processing pause and eventual removal. org's own The price of low-signal supply is becoming lost access, not merely a smaller check.

Section 6Counterchecks

Congestion is real, but triage may automate too

Contrast reports 5 percent agreement among three AI scanners and 17 percent repeatability for one scanner across runs. It estimates $315 thin in model-token cost to scan two million lines and $128,000 thin to triage the output. The accessible methods are incomplete and the vendor sells the alternative it recommends, so the figures are directional, not dispositive.

A Splunk practitioner reports product-aware triage up to 20 times faster while retaining evidence gates and human review. thin It is a self-report, not an external benchmark, but it prevents a simple "cheap discovery means permanently expensive triage" conclusion.

Discovery volume is not exploitation value

VulnCheck found 14 confirmed exploited vulnerabilities among 1,061 attributed to AI-assisted discovery: 1.3 percent, roughly the overall first-half rate. single Of more than 23,000 Anthropic Project Glasswing findings, it found 126 published CVEs and one confirmed exploited vulnerability. This is the strongest available counter-evidence to treating discovery volume as a direct proxy for offensive value.

Section 7July falsifier scorecard

July falsifierStatusWhy
Broker publishes an AI/LLM category with a priceNot triggeredNeither checked broker added one
Broker cuts chain prices because of automation, or a realised sale clears below prior anchorsNot triggeredPublic maxima unchanged; no credible realised transaction located
Public system reproduces withheld hardened-target resultsNear missAstra is gated and vendor-run; Hugging Face is real but a different target class
Apple or Google cuts top-chain rewardsNot triggeredThe top anchors remain
A cut program proves triage cost made payouts bindingNot triggeredContrast is not a cut program and methods are incomplete
GitHub's post-restructure signal remains unchangedUnresolvedNo public outcome series

Section 8Four rival explanations, at full strength

Hypothesis 1
The technical capability wall moved.

Strong for multi-system logic and configuration chains. For hardened browser and OS chains, the support is still vendor-run and not independently reproduced.

Hypothesis 2
A broad price collapse is underway.

Supported in parts of the middle tier. Contradicted by unchanged top anchors and rising aggregate bounty payments. Private realised broker prices remain opaque.

Hypothesis 3 · best fit
Access is now the binding scarcity.

OpenAI, Google and Anthropic independently converged on trusted or restricted cyber access while generally available discovery products became cheaper.

Hypothesis 4
This is mostly queue congestion, not capability economics.

Explains access sanctions, submission limits and some middle-tier compression. It is incomplete because triage itself is automating and the strongest systems remain deliberately scarce.

The machine discount now applies to more than defects and primitives, but not uniformly. Multi-system chaining is real, hardened-target chaining is credibly claimed, and both remain separated from the public market by access controls, evaluation opacity and target class.

Section 9What would change this view next

  1. Astra's system card, wider availability and an independent hardened-target reproduction.
  2. A public broker AI category, a changed chain maximum or a documented realised transaction.
  3. GitHub post-restructure validity, volume, response-time and payout data.
  4. METR's review of Anthropic's real-system incidents.
  5. Independent replication of Google's Flash Cyber discovery and patching claims.
  6. Program-scale evidence that product-aware triage reduces queue cost.
  7. Apple data on how processing pauses affect valid submissions and response time.

Section 10Sources and method

This point-in-time pass audited 27 saved sources. The machine-readable snapshot carries the evidence class and caveat for every load-bearing row.

Full synthesis: research/2026-09-02-gated-capability/RESEARCH.md. Structured observations: data/september-2026-snapshot.json.

Capability and incidents

Model price and access

Prices, policy and market controls

Limits

This is a 2 September snapshot, not a September retrospective. Public maxima are not realised prices. Vendor benchmark claims remain labeled vendor claims. Absence of a public transaction is not proof that a private market did not move.