$_ stdout

Patchmageddon? Or Not So Much

J.P. Morgan's thirty-page report says AI has collapsed the gap between disclosure and exploitation from a year to a day. The data behind that claim is solid. The conclusion most readers will draw from it — patch everything, faster — is not, and the breach data says why: this is a prioritization problem, not a speed problem.

The exploitability gap and the breach data throughout this post say something less alarming: fewer than 2% of disclosed vulnerabilities have ever driven real-world risk, and that concentration hasn't budged even as raw discovery volume exploded.

The fix isn't patching faster across the board — it's prioritization: pointing that speed at the vital few flaws that are actually reachable and actually exploited.

BL Dr. Ben Livshits August 4, 2026 · 58 commits
Cover page of J.P. Morgan's Eye on the Market report, 'Patchmageddon: The race to patch software vulnerabilities before zero-day cyber-exploitations proliferate,' July 2026, by Michael Cembalest — illustration of four cloaked riders on horseback overlooking a lit city skyline at night.

When J.P. Morgan publishes a thirty-page research report titled Patchmageddon, the security industry takes notice. Written by Michael Cembalest, Chairman of Market and Investment Strategy, alongside Global CISO Pat Opet and the bank's cybersecurity leadership, the report formalizes a thesis that has been building for a year: frontier AI is industrializing vulnerability discovery, collapsing exploitation timelines, and driving CVE growth to levels the old disclosure-and-patch cadence was never built to absorb.

I've been tracking the same shift from a different angle — the median-time-to-exploit curve sliding from roughly a year in 2021 down toward zero, and the bug-bounty pipeline drowning in AI-generated noise as a result. J.P. Morgan's report covers similar ground with a bank's resources behind it, and it's worth taking seriously. But it also invites a question its own data doesn't answer: does an explosion in disclosed vulnerabilities equate to existential enterprise risk? Pair the report's warnings with the breach data that's since come out, and the picture gets both scarier and more tractable at the same time.

Vulnerability discovery and exploit generation now run at machine speed — AI systems finding, chaining, and weaponizing flaws in hours. Enterprise remediation still runs at human speed — change control, regression testing, maintenance windows measured in months. The gap between those two clocks is the whole story.
Infographic titled 'Patchmageddon: The AI-Driven Race for Cybersecurity Survival.' Left: a tornado graphic under the heading 'The Crisis: A Narrowing Window for Defense,' noting AI-driven vulnerability discovery outpacing patching by roughly 16.5 times and that close to 80% of exploits are now zero-day, striking on or before public disclosure. Center: a line chart, 'The Speed Gap,' showing attacker time-to-exploit falling from about one year in 2021 toward under a day by 2027, against organization time-to-remediate climbing over the same period. Right: 'The Solution: Fighting Fire with AI' — deploying defensive AI agents like Claude Security and GPT-5.5-Cyber, pivoting evaluation toward time-to-patch, and hardening unpatchable legacy hardware with network segmentation.
// Figure 1. The crisis-and-solution framing in one graphic — discovery outpacing patching, zero-day-dominant exploitation, and the speed gap this whole post is about. The response panel's "patch faster" framing is the part this post pushes back on.

For twenty years, defenders operated on the assumption that vulnerability research was a human-constrained bottleneck: someone had to read the code, understand the bug class, and hand-craft an exploit. That assumption is dead. The interesting questions aren't whether the gap above is real, but what to do about it — and whether the disclosed-volume panic the report invites is even pointing at the right thing.

01 The Hard Data: Collapsing Timelines and the Patch Gap

Five numbers carry most of the argument's weight here, and they're worth isolating before turning to the charts that back them up.

The Five Numbers

The empirical core of J.P. Morgan's report is a widening divergence between attacker capability and defensive response:

Two Clocks, Diverging

Put the two trend lines side by side and the shape of the problem is obvious — one curve heading toward zero, the other drifting upward:

Attacker time-to-exploit versus enterprise mean-time-to-remediate, 2020–2026 Line chart with two series from 2020 to an estimated 2026. Attacker time-to-exploit falls from about 30 days in 2020 to 18 days in 2023 to 1 day in 2025 to under zero, meaning pre-disclosure exploitation, by 2026. Enterprise mean time to remediate rises from about 60 days in 2020 to 58 days in 2023 to 88 days in 2025 to more than 95 days by 2026, moving in the opposite direction across the same window. Attacker TTE vs. enterprise MTTR, 2020–2026 0d MTTR ~60d ~58d ~88d ~95+d TTE ~30d ~18d ~1d <0d (pre-disclosure) 2020 2023 2025 2026E
// Figure 2. Attacker time-to-exploit collapsing toward — and past — zero, while enterprise mean-time-to-remediate keeps climbing in the opposite direction: the two clocks whose gap this whole post is about. Source: J.P. Morgan, Patchmageddon (2026), citing Sidero Labs (May 2026).

Rate vs. Speed

None of this comes from one source. J.P. Morgan's report synthesizes work from Mitiga, Forescout, Sonatype, the UK AI Security Institute, independent researchers Sergej Epp — creator of the Zero Day Clock dashboard — and Josh Bressers, plus model evaluations published directly by Anthropic and OpenAI. That's a wide enough base of independent instrumentation that the top-line trend is hard to argue with, whatever you think of any single number in it.

Two charts: total CVEs published per year climbing from under 20,000 in 2018 toward 50,000 by 2025 while their exploit rate first rises then falls sharply toward 2026; and mean and median time-to-exploitation both collapsing from roughly 800 and 750 days in 2018 to near zero by 2026
// Figure 3. Total CVEs published and their exploit rate (left), against mean and median time-to-exploitation (right) — both clocks bending toward zero at once. Source: Zero Day Clock, July 2026.

Two things are happening in that chart at once. The left panel's CVE bars climb every year while the exploit-rate line — the share of that growing pile ever weaponized — actually peaks around 2021 and falls after, because disclosure volume is outrunning exploitation volume. The right panel is the sharper story: mean and median time-to-exploitation both collapse from the better part of two years to near zero across the same window — the number J.P. Morgan's report leans on hardest. Rate and speed are answering different questions, and conflating them is most of what goes wrong in the "48,000 CVEs" framing.

More Things, Faster

Figure 4 tells the volume side of the same story from a different angle: weaponized exploitations — confirmed attacks, not just published CVEs — more than doubled between 2018 and 2024, even as the TTE lines above keep falling toward zero. Attackers aren't exploiting a smaller number of things faster; they're exploiting more things, faster, at the same time. That combination is what makes the remediation-drag bullet above bite harder than the disclosure-surge bullet does on its own.

Bar and line chart showing weaponized exploitations per year rising from around 200 in 2018 to over 500 by 2024, overlaid with mean and median time-to-exploitation lines both falling from roughly 800 days to near zero by 2026
// Figure 4. Weaponized exploitations by year against the same time-to-exploitation collapse — volume climbing while the response clock keeps falling. Source: Zero Day Clock, July 2026.

The Survival Curve

The two panels in Figure 5 pull in opposite directions, and both matter. The left one is the source of the 80%-and-climbing zero-day figure above: exploitation increasingly lands before or on disclosure day, not after. The right one behaves like a survival curve — for each cohort year, it plots what share of that year's CVEs have been exploited by a given number of days out, and every cohort levels off well under 100%, including the older ones that have had years to accumulate exploits.

Two charts: zero-day exploitation rate climbing from under 20% in 2018 to nearly 80% by 2026, showing exploits increasingly landing on or before disclosure day; and the share of each year's CVEs eventually exploited after disclosure, by cohort, with newer cohorts (2025, 2026) climbing far more slowly than older ones
// Figure 5. The zero-day-dominant share of exploitation (left) — the 80% figure above — and, cohort by cohort, how much of a given year's CVEs ever get exploited at all as days pass since disclosure (right). Source: Zero Day Clock, July 2026.

Read the newest cohorts carefully, though: 2025 and 2026 sit far below the rest mostly because they haven't had time to run yet, not necessarily because they're safer. It's the same shrinking-share-of-a-growing-pile pattern this post returns to under Understanding What's Exploited.

02 Frontier AI and Machine-Speed Discovery

The collapsing timelines in the last section have a specific cause: a new generation of models built or tuned specifically for security work, and the pipeline problems that follow once they start finding bugs faster than anyone can triage them.

The Models Doing the Work

The driver behind the shift is a new class of frontier models purpose-built or fine-tuned for security research — Anthropic's Claude Mythos Preview (built on the Fable 5 line) and OpenAI's GPT-5.5-Cyber, paired with tooling like Codex's security plugins. These aren't general-purpose assistants that happen to be decent at CTF problems; they're systems specifically evaluated on vulnerability discovery, and the numbers reflect it.

Frontier scale isn't the only lever, though. Cisco's 3-billion-parameter Antares matches frontier-level vulnerability-localization accuracy for under $1 a run versus roughly $141 for a leading proprietary model, and Google's cyber-specialized Gemini 3.5 Flash Cyber — built on Google's smallest, cheapest tier — found more V8 vulnerabilities than Claude Opus 4.6 did. A blog post on using small language models (SLMs) makes the fuller case: the durable edge in AI-driven discovery is the orchestration system wrapped around a model, not the model's parameter count.

The Triage Bottleneck

// the triage funnel 23,000 candidate findings taken to review. 3,900 deemed high- or critical-severity. 530 formally reported to maintainers. 75 patched. Tuskira measured discovery outpacing patching by 16.5× — even with maintainers acknowledging 90% of the reports they received.

That last point deserves more weight than the report gives it. A frontier model that finds real bugs faster than humans is a capability question. A frontier model that finds real bugs faster than a single unpaid maintainer can triage them is a staffing question, and staffing questions don't get solved by better AI on the discovery side — they get solved by AI, or something, on the triage side.

The Defensive Response

The Patchmageddon report spends its remediation chapter on exactly this: Claude Security patched 2,100 vulnerabilities in the weeks after launch, and OpenAI's Codex Security has scanned more than 30 million submissions across 30,000 codebases, with 70,000 discoveries marked repaired by human reviewers and another 500,000 auto-determined as repaired.

That is not a footnote to the panic thesis — it is the only answer shape that scales, and the report is more useful for putting it in writing than for the headline.

That's one pillar of a broader defensive architecture. A recent post traces the full four-pillar stack the defensive literature has converged on — harden, watch, contain, patch — along with the caution that belongs next to any autonomous-remediation number: stripped of a hint about where a bug lives, frontier models patch real zero-days correctly on fewer than 15% of cases; handed the location and nature up front, the best of them jumps to 95.7%. The gap between those two numbers is most of what "autonomous remediation" still has to prove.

03 Why Counting CVEs Isn't the Answer

The case against volume-driven panic starts with a simple observation about what a raw count actually measures.

The Vanity-Metric Problem

Here's where I think the report, read uncritically, leads people astray. Treating the raw explosion in CVE counts as an existential threat to corporate survival misunderstands how operational risk actually works. As security analyst Adrian Sanabria has pointed out, raw vulnerability counts have never been a reliable predictor of organizational failure — they're a count of things that could go wrong, not things that do.

Strikingly, the Patchmageddon report knows this. A footnote in its state-of-the-world chapter acknowledges that of the rising tide of reported vulnerabilities, "1.5%–2.0% get exploited each year."

That's the same figure this section arrives at from independent sources — a number so obvious the report itself has to stipulate it in passing, then bury it under a thirty-page case for exactly the kind of volume-driven panic that its own stipulation undercuts. The post isn't catching J.P. Morgan omitting the breach data; it's catching the report acknowledging it and pressing on anyway.

Vulnerable Does Not Imply Exploited

The academic grounding for this goes back further than the current panic. Daniel Perez and I surveyed 23,327 Ethereum contracts flagged as vulnerable by six academic detectors in Smart Contract Vulnerabilities: Vulnerable Does Not Imply Exploited (30th USENIX Security Symposium, 2021).

Despite the potential risk and financial advantages of exploitation, only 1.98% had ever actually been exploited — at most 8,487 ETH (~$1.7 million) out of the 3 million ETH (~$600 million) those detectors said was at risk, or 0.27% of the money supposedly on the table.

The gap wasn't random: the funds actually at risk were concentrated in a handful of contracts, while the rest of the flagged population was vulnerable on paper and functionally unexploitable — the same reachability gap the report cites today, three years before "reachability" was the word anyone used for it.

Table from Perez and Livshits (2021), Figure 11, titled 'Understanding the exploitation of potentially vulnerable contracts.' Six vulnerability classes (RE, UE, LE, TO, IO, UA) broken down by vulnerable contracts, total Ether at stake, transactions analyzed, contracts exploited, percent of contracts exploited, exploited Ether, and percent of Ether exploited. Totals: 23,327 vulnerable contracts, 3,124,433 Ether at stake, 463 contracts exploited (1.98%), 8,487 Ether exploited (0.27%).
// Figure 6. Figure 11 from the original paper, reproduced directly: the same six-vulnerability-class breakdown behind the 1.98%-of-contracts, 0.27%-of-Ether numbers above. RE = re-entrancy, UE = unhandled exceptions, LE = locked Ether, TO = transaction order dependency, IO = integer overflow, UA = unrestricted action. Source: Perez & Livshits (2021).

What Matters for Exploitability

Four mechanisms explain the gap:

// the reachability filter Static code base → contains a vulnerable library. Runtime execution → is the vulnerable function actually reachable, ideally confirmed with a working proof-of-concept exploit rather than just a runtime trace? Perimeter controls → is the service exposed externally? Only when every gate answers "yes" does a CVE translate into real operational risk — and for most CVEs, at least one gate answers "no."

The finding held up under a much bigger microscope a few years later. Hu et al. re-ran the question in 2024 across 136,969 real-world contracts — six times the scale of the original study — using seven separate detectors. Of the 4,364 contracts flagged as vulnerable, 75.25% turned out to be unexploitable outright: false positives or code with no real security risk once you looked closer. Of the 1,080 that were genuinely exploitable, only 66 — 6% — were ever actually exploited.

Run the full funnel and roughly 1.5% of everything flagged as vulnerable ever got touched — a number that will look familiar by the time this post gets to VulnCheck and EPSS in the next section, and to Root Evidence's breach data two sections after that.

Smart contracts and enterprise software are about as different as codebases get, and the same shape of answer keeps showing up.

Exploits to Dollars

A third paper adds the missing piece: if presence doesn't equal exploitation, what actually turns a flagged vulnerability into a real loss?

Rezaei, Eshghie, et al.'s 2025 SoK stays in the same domain as Perez and Livshits — smart contracts, where the presence-versus-exploitation tension is sharpest of all, because an exploit isn't a theoretical risk score, it's money leaving the contract in real time. They traced 50 real attacks from 2022 to 2025 that cost more than $1.09 billion combined, against a catalog of 24 active vulnerability classes drawn from 71 academic papers. Most of that billion dollars didn't come from a single isolated bug — it came from exploit chains: protocol-logic flaws, governance and lifecycle issues, and external dependencies combining with an implementation bug, rarely any one of those alone.

Academic vulnerability research, the authors note, still mostly studies implementation bugs — the last item on that list. Protocol-logic flaws, governance and lifecycle issues, and external dependencies — the other three — are where the money actually goes.

04 Understanding What's Exploited

The Perez-and-Livshits result says presence and exploitation are different distributions. It stops short of saying just how different — for that you need actual exploitation data, gathered more than one way. Two more sources, using two different methodologies, land close enough to the same number that it stops looking like an artifact of any one analyst's assumptions.

Two Convergent Measurements

Two independent measurement approaches converging on the same single-digit exploitation rate Diagram showing VulnCheck's direct exploit-tracking census finding that 1 percent of 2025-disclosed CVEs were ever confirmed exploited, and Picus Security's EPSS-based analysis finding that 2.3 percent of CVSS 7-plus CVEs have ever been observed under exploitation. Two unrelated methodologies, neither built on the other, converging on the same low single-digit percentage. Two methods, one number VulnCheck direct exploit-tracking census EPSS machine-learned probability model ~1–2.3% of disclosed CVEs ever confirmed or observed exploited Neither team built on the other's work
// Figure 7. Two unrelated methodologies — a direct exploit census and a machine-learned probability model — landing on the same low single-digit exploitation rate. Sources: VulnCheck (2026); Picus Security (2026).

Two different measurement approaches — one an exploit-tracking census, one a machine-learned probability model — converging on a single-digit percentage is about as strong as evidence gets in this argument: not one team's methodology, but the same signal recovered independently.

// two methods, one number VulnCheck's direct census: 1% of everything disclosed in 2025 was ever confirmed exploited. Picus Security's EPSS-based analysis: 2.3% of CVSS 7+ CVEs have ever been observed under exploitation. Neither team built on the other's work.

The Prioritization Triad

If raw volume is a distraction and discovery keeps accelerating, the honest answer is to stop measuring the wrong thing. The shift is from volume-based remediation to evidence-based risk reduction, and it rests on three signals working together, not any one of them alone:

Threat intelligence and breach data (is this being exploited, right now, anywhere) meets organizational context (does this asset matter, and is it reachable from outside) meets dynamic reachability (can you build a working proof-of-concept against my deployment, not just trace that the code path executes). A CVE that fails any one of the three doesn't get to jump the queue just because its CVSS score is high.

  1. Threat intelligence and breach data. Prioritize vulnerabilities with demonstrated, active in-the-wild exploitation — CISA's KEV catalog and breach tracking — over theoretical CVSS predictions.
  2. Organizational context and exposure. Weigh asset value, network position, and internet reachability before assigning priority. An unpatched edge VPN is immediate exposure; an isolated internal utility library is not.
  3. Dynamic reachability analysis. Confirm that a vulnerable code path actually executes in your specific stack — runtime tracing, taint analysis, and sandboxed testing all count as evidence, and a working proof-of-concept exploit is simply the strongest form that evidence can take — before spending engineering time on it.

The Limits of Threat Intelligence Alone

That first leg, alone, isn't enough — EPSS is the clearest illustration of why. Splunk's Muhammad Raza credits its breadth, more than a thousand input variables, with letting it outperform CVSS at separating signal from noise. The score is meant to be read literally, not just ranked: ReversingLabs' John P. Mello Jr. explains that a 90% EPSS score means roughly nine in ten CVEs rated that way actually get exploited, a calibration claim EPSS's own maintainers back with a recommended action threshold of 0.36; below that, a CVE isn't supposed to jump the queue.

Independent researcher Rianna Parla tested that threshold against reality.

Of 250 CVEs added to CISA's KEV catalog during her study period, more than two-thirds sat below 0.36 the day before inclusion — the opposite of what the threshold is supposed to catch.

Narrowed to the 57 most recent entries, 42 carried a consistently low, mostly single-digit score right up to confirmed exploitation; four were added to KEV the same day they were disclosed, already under attack before EPSS had a chance to say anything at all.

Share of KEV-catalogued CVEs above and below EPSS's own prioritization threshold before inclusion Of 250 CVEs added to the CISA KEV catalog during the study period, roughly a third scored at or above EPSS's own recommended 0.36 prioritization threshold just before inclusion. More than two-thirds scored below that threshold, meaning a defender following EPSS's own cutoff would have left most of them unprioritized until after exploitation was already confirmed. 250 KEV CVEs: EPSS score just before inclusion ~1/3 more than 2/3 at or above 0.36 EPSS's own cutoff below 0.36 unprioritized until after the fact Source: Parla (2024), "Efficacy of EPSS in High Severity CVEs found in KEV"
// Figure 8. EPSS's own recommended threshold, tested against reality: most confirmed-exploited CVEs sat below the cutoff meant to flag them.

One case makes the trailing-indicator problem concrete on its own: CVE-2020-17519, an Apache Flink path-traversal flaw, has carried a 90% EPSS score continuously since the model's third version launched in March 2023 — a confident, correct-looking call, until Palo Alto Networks' Unit 42 incident data places the real exploitation window at November 2020 to January 2021, more than two years before the model that's "predicting" it even existed.

CVE-2020-17519 (Apache Flink): EPSS score versus actual exploitation timeline Timeline showing actual in-the-wild exploitation occurred November 2020 to January 2021, per Palo Alto Networks Unit 42. EPSS version 3 launched in March 2023 and scored this CVE 90 percent from day one. The CVE was added to the CISA KEV catalog in May 2024. The score remains around 90 percent today. The gap between real exploitation and the model even existing is more than two years. CVE-2020-17519 (Apache Flink): score vs. reality Nov 2020–Jan 2021 Actual exploitation Mar 2023 EPSS v3 scores it 90% May 2024 Added to CISA KEV Present Still ~90% 2+ years before the model existed
// Figure 9. A 90% EPSS score that's correct today and would have been useless in 2020: the model didn't exist yet when the actual exploitation happened. Source: Parla (2024), citing Palo Alto Networks Unit 42.

Parla's wider sample shows the same shape: CVEs from 2022 or earlier almost universally score high before their KEV listing, while CVEs from 2023 onward — the ones a defender would actually need a forecast for — turn in "mixed-to-poor" results. A model that looks sharp in hindsight and blurry in the moment it matters is describing the past, not forecasting the future.

Global Signal, Local Risk

The other gap is exactly the triad's second leg: EPSS scores the whole internet, not your environment. Kodem's Mahesh Babu makes the point with a concrete case: CVE-2024-34341, a cross-site-scripting flaw in the Trix editor, scored around 20% — comfortably below any reasonable cutoff — yet was a critical risk in internet-facing deployments with privileged users, a distinction EPSS's global average has no way to see. His framing: EPSS is a weather forecast for the whole region, not a report on whether your street floods.

None of this makes EPSS worthless. It's a genuine improvement on CVSS's static severity score, and the IsMalicious team notes its daily update cycle catches newly weaponized CVEs within 24 hours of exploit code surfacing. It's simply not the standalone predictor its name promises — which is the whole argument for treating it as one leg of three, not the entire strategy.

Root Evidence's own Robert "RSnake" Hansen reaches the identical conclusion from the other direction, and cites FIRST's own guidance to back it: EPSS "is best used when there is no other evidence of active exploitation," and where such evidence exists, it "should supersede the EPSS estimate." Hansen's target is the reverse ordering some teams propose — EPSS first, KEV lists second — which he points out has no coherent stopping rule once you notice EPSS and KEV both cover the same 300,000-plus CVEs. Evidence of actual exploitation, not a probability estimate of it, is supposed to lead. That's the same argument from earlier in this post, restated by a vendor that sells the opposite of what EPSS sells.

// second, not first FIRST's own guidance ranks EPSS behind actual evidence of exploitation, not ahead of it. A KEV listing or a breach report outranks a probability score every time — which is exactly what "one leg of three" means in practice.

Speed still matters — the report isn't wrong about that. But a 24-hour remediation cycle aimed at the 1.4% of vulnerabilities carrying real exploitable risk protects an organization far more than 90 days spent clearing thousands of flaws nothing can ever reach.

05 What Actually Mattered in Q1 2026

A third, independent data source makes the same case again, this time from breach outcomes rather than theory.

The Q1 2026 Report

For a third cut at the same question — this one built directly on breach data rather than tracked exploits or predicted probabilities — look at Root Evidence's Q1 2026 analysis, Stop Counting CVEs: What Actually Mattered in Q1 2026.

Founded by longtime security researchers Robert Hansen, Jeremiah Grossman, Heather Konold, and Lex Arquette, Root Evidence ran publicly available vulnerability data through a real-world breach lens instead of a severity score.

// NVD CVE severity vs. real-world exploitation
Category Count Share
Total analyzed CVEs352,357100.0%
Known exploited vulnerabilities4,9201.4%
Unexploited in the wild347,43798.6%

The findings cut hard against volume-driven panic:

Under 0.2%

Root Evidence followed the Q1 report with a companion post, "Why We Built the Evidence Platform," that cuts the same data even harder. Running cyber insurance claims and incident-response data against the full CVE population — more than 370,000 published, 40,000 to 60,000 new every year — turned up fewer than 600 CVEs ever tied to an actual financial loss.

That's under 0.2%, an order of magnitude tighter than the 1.4% exploited figure above, because it asks a stricter question: not just exploited, but exploited enough to cost someone money. You would never go to Vegas with those odds.

Betting on It: The Mythos Warranty

The company backs that number with a bet of its own: the Mythos Warranty — unrelated to Anthropic's Claude Mythos despite the shared name — reimburses customers up to $5 million for losses tied to any CVE Root Evidence failed to flag, underwritten by cyber insurers who reviewed the same breach data. A vendor's confidence in a prioritization method means more when it's backed by real money than by a press release.

Cyber insurance is the right place to look for this kind of signal, because the market has already been forced through a version of the same reckoning this post describes.

The Insurance Market's Own Blind Spot

It's a roughly $23 billion line of business growing near 15% a year, and underwriters have been moving away from annual paper questionnaires toward continuous telemetry: Coalition and At-Bay both price and monitor policies off live signals — external attack-surface scans, MFA and EDR deployment status, and increasingly the kind of exploitation data cited throughout this post.

But most of that market still leans on the same blunt instrument this post has argued against throughout. A large share of cyber policies now carry a "known vulnerability" or failure-to-patch exclusion: if a breach traces back to a CVE that was publicly disclosed and patchable months earlier, the insurer can decline the claim outright.

Coalition's own analysis put a number on how blunt: as of mid-2025, more than 61,000 disclosed vulnerabilities were broad enough to trigger that exclusion, while barely 1% of them sat in CISA's KEV catalog of confirmed exploited flaws.
Over 40% of cyber claims now get denied, and exclusions like this one are the leading reason. The insurance industry, in other words, has its own CVE-counting problem — just measured in payouts instead of remediation hours.

That's what makes the Mythos Warranty notable beyond the marketing. Standard exclusions penalize a customer for any unpatched CVE that existed, reachable or not.

Root Evidence's warranty inverts that logic: it pays out specifically when a CVE it failed to flag as high-risk turns out to have been reachable after all. One model insures against the presence of vulnerabilities; the other insures against a failure of prioritization.

06 Business Imperatives and Policy Solutions

Closing the gap between the two clocks takes action on both the corporate and the government side — but the throughline in every recommendation below is still prioritization, not raw speed. J.P. Morgan's report is more useful here than in its headline framing — the recommendations are concrete.

Enterprise Action Items

A third lever works differently from the first two: predictive triage. KEV and breach data both look backward — something has to have already happened before either can flag it. Suciu et al.'s Expected Exploitability research (see Related Work) shows a calibrated model predicting whether a functional exploit will ever get built, rather than waiting for one to show up, can lift precision from 49% to 86% over baseline classifiers. Feed a model like that into the threat-intelligence leg of the triad above as a forward-looking supplement to prioritization — not as a replacement for confirmed exploitation evidence or reachability.

The fourth lever, for legacy operational technology (OT), is the hardest case: OT defense. It can't always take a patch — up to 30% of it, running on 10-to-18-year hardware, can't accept one at all — so compensating controls (segmentation, monitoring, anomaly detection) have to stand in instead. The catch, below, is that most operators can't see their network well enough for those controls to work.

// the visibility gap The Dragos 2026 Operational Technology Cybersecurity Year in Review found that 30% of incident-response cases in 2025 started not with a detected intrusion or a ransom note but with someone noticing "something seemed wrong," and in the majority of those cases the data needed to confirm whether cyber was involved had never been collected. OT telemetry is transient; if it isn't recorded when it happens, it's gone. 90% of Dragos's asset-owner clients still can't detect the decade-old ELECTRUM technique that just resurfaced in Poland.

Government and Policy Mechanics

Because so much critical infrastructure is privately owned, neither lever — government disclosure programs, or enterprise action on them — works alone. Enterprises without a clearinghouse feed are triaging blind; a clearinghouse with no enterprise capacity to act on it just moves the bottleneck upstream.

07 Related Work

The presence-versus-exploitation gap this post keeps returning to has its own research lineage, stretching back well before EPSS existed. Six papers, spanning several different measurement approaches, are worth setting down explicitly — in roughly the order the field tackled them.

Zero-days, measured in the field. Bilge and Dumitraș's "Before We Knew It" (ACM CCS 2012) is the paper that put a number on how long a zero-day actually survives before disclosure: field telemetry from 11 million real hosts turned up 18 true zero-day attacks, running a median of 8 months and an average of roughly 10 months before anyone noticed — a much slower clock than the "under a day" figures this post's own charts are built on, and a reminder that the collapsing-TTE story is specifically about disclosed, not undiscovered, vulnerabilities.

Twitter as a leading indicator. Sabottke, Suciu, and Dumitraș's "Vulnerability Disclosure in the Age of Social Media" (USENIX Security 2015) predates EPSS by five years and reaches a similar structural conclusion from a completely different signal: a CVSS-only baseline classifier tops out under 9% precision, and folding in Twitter chatter about a CVE measurably improves on it — an early demonstration that the static severity score alone was never going to be enough.

The original EPSS paper. Jacobs, Romanosky, Adjerid, and Baker's "Improving Vulnerability Remediation Through Better Exploit Prediction" (Journal of Cybersecurity, 2020) is the methodology this entire post's EPSS discussion assumes but never names directly: 75,000+ CVEs from 2009–2018, real exploit observations from roughly 100,000 monitored corporate networks, and a gradient-boosted model that reaches 70% coverage of exploited vulnerabilities by patching about 7,900 CVEs — versus 30,000 to 34,000 under a CVSS- or exploit-count-based strategy.

The same author, seven years later. Suciu, Nelson, Lyu, and Bao's "Expected Exploitability" (USENIX Security 2022) continues the line the Twitter paper started, now predicting whether a functional exploit will ever get built, not just whether one shows up on social media. Across 103,137 vulnerabilities, their metric lifts precision from 49% to 86% over prior classifiers — and the paper opens with CVE-2017-0144, the flaw WannaCry and NotPetya weaponized, sitting outside every expert patch-priority list right up until it wasn't.

OT/ICS gets its own scoring model. Yoon, Kim, Kim, and Euom's "Vulnerability Exploitation Risk Assessment Based on Offensive Security Approach" (Applied Sciences, 2023) makes the case this post's own OT material only gestures at: standard vulnerability assessment doesn't transfer to industrial control systems, so the authors build a dedicated OT/ICS risk score from exploit-chain risk, exploit-code availability, and exploit-use probability — for exactly the kind of environments Dragos tracks earlier in this post.

The weakness-type version of the same question. Mell, Bojanova, and Galhardo's "Measuring the Exploitation of Weaknesses in the Wild" (2024) runs the presence-versus-exploitation test one level up, at the CWE weakness-class level rather than the individual CVE: across 130 weakness types tracked from April 2021 to March 2024, 92% were not being constantly exploited — the same shape of answer this post keeps finding, recovered again at a different unit of analysis.

Timeline of presence-versus-exploitation research, 2012 to 2024 Six papers plotted chronologically: Bilge and Dumitras, 2012, zero-day field study; Sabottke et al., 2015, Twitter signal; Jacobs et al., 2020, the original EPSS paper; Suciu et al., 2022, functional exploit prediction; Yoon et al., 2023, OT/ICS scoring; Mell et al., 2024, weakness-class level analysis. Presence-versus-exploitation research, 2012–2024 Zero-day field study Bilge & Dumitraș 2012 2015 Sabottke et al. Twitter signal EPSS itself Jacobs et al. 2020 2022 Suciu et al. Functional exploits OT/ICS scoring Yoon et al. 2023 2024 Mell et al. Weakness-class
// Figure 10. The same six papers, plotted: a twelve-year run from field-measured zero-days to a weakness-class refinement of the exact question this post keeps asking.
Read together, they cut against the same instinct from six different directions and three different eras: a static severity score was never going to predict exploitation, whether the signal added instead is social media chatter, field telemetry, a machine-learned model refined over a decade, or a domain-specific score built for hardware that can't be patched at all.

08 Open Questions

The argument above is that disclosed-volume panic overstates enterprise risk — not that the underlying capability is staying confined. The Patchmageddon report spends five pages on the part of its thesis this post has given the least weight, and that part is the one that holds up best.

Epoch AI estimates freely downloadable open-weight models now lag the leading proprietary frontier by roughly three months, not years; Kimi K3 has already reached cyber parity with Anthropic's and OpenAI's frontiers in J.P. Morgan's own testing, at roughly half the per-task cost. And once a model of that class is downloadable, guardrails are removable — a tool called Heretic strips safety protections from open-weight systems in under ten minutes on a laptop, and a jailbroken Z.ai GLM 5.2 has been circulating on Hugging Face since June.

This critique says most CVEs don't matter; it does not say capability won't democratize. Both can be true, and the second is the reason prioritization matters more, not less.

With that on the table, the picture above is solid on where the field has been and less settled on where the actual boundary of AI-driven offense and defense sits right now — and that boundary is mostly being tested in working papers, not asserted in press releases — most of them sitting on arXiv rather than in a peer-reviewed venue yet. Seven recent ones are worth flagging.

RQ1: Exploitation Capability

ExploitGym (Wang, Schiller, Li, et al., 2026) benchmarks AI agents against 898 real vulnerability instances instead of counting CVEs, and even frontier models fall short of the machine-speed narrative once forced to produce a working exploit: Claude Mythos Preview and GPT-5.5 succeeded on 157 and 120 instances respectively — a "non-trivial fraction," in the authors' words, not a majority.

Is that ~18% success rate a ceiling imposed by something structural in the exploit-construction task, or a snapshot of where the frontier happens to sit today — the 2024 model evaluated against a 2026 benchmark, with the next generation already training?

RQ2: Triage Economics

OpenAnt (Korda & Evron, 2026) operationalizes the reachability filter from earlier in this post directly: it decomposes a codebase into units filtered by reachability from external entry points before running deeper analysis, cutting the search surface by up to 97%. The largest repository runs in the paper still cost several hundred dollars each.

Does that cost curve bend down fast enough to matter at ecosystem scale, or does reachability-gated triage stay a boutique technique?

RQ3: Slop Sustainability

A July 2026 study of GitHub contribution data — "AI Slop is DDoSing Open Source" (Afroz, Miller, Menezes, Gilmour, Sarma & Feng) — found pull-request volume rising through 2025 while merge rates fell, with first-time contributors seeing an 18% drop against projections: the maintainer side of the triage-overload problem raised in Frontier AI and Machine-Speed Discovery.

Do better filtering tools arrest the merge-rate collapse, or does volunteer maintenance simply not survive contribution at this volume?

RQ4: Fabricated Exploits in the Wild

VulnCheck's 2026 Exploit Intelligence Report documents AI-generated PoC repositories that look functional but never exercise the claimed vulnerability — including a fake "working" React2Shell exploit that cost defenders real investigation time, and a fabricated CVE attribution that Google's AI search summaries then repeated as fact, citing the fabricated repository as its source.

When one AI system's fabrication becomes another's ground truth, does confirmation need more automation — or a human, specifically because the automation is what's being fooled?

RQ5: The Confirmation Bottleneck

If discovery is automated and reachability filtering helps, the next chokepoint is confirmation — something still has to verify a flagged report is a real, working exploit rather than a static-analysis false positive. AXE (Sajadi, Nguyen, Damevski & Chatterjee, 2026) automates that step and gets a working proof-of-concept on 30% of its benchmark — three times better than prior black-box approaches, and still well under a third.

Can automated confirmation keep pace with automated discovery, or does it just relocate the bottleneck?

RQ6: Benchmark Validity

CAIBench (Sanz-Gómez, Mayoral-Vilches, Balassone, et al., 2026), a meta-benchmark spanning more than 10,000 offensive and defensive test instances, found models saturating on cybersecurity knowledge questions (~70%) while dropping to 20–40% on multi-step adversarial scenarios. Knowing what a vulnerability is and being able to exploit or defend one are different capabilities — which raises a question about every "frontier model found N zero-days" headline, including the ones cited earlier here.

Do those "N zero-days found" headlines measure actual offensive capability, or only textbook recall?

RQ7: Small Models, System Edge

The Antares and Gemini 3.5 Flash Cyber results cited earlier point at a structural question that none of the benchmarks above resolves: if a 3-billion-parameter model matches frontier-level localization accuracy at 1/141st the cost, and the durable edge is the orchestration system wrapped around the model rather than the model itself — as a recent post on small language models argues — then the frontier's advantage narrows to the slice of the task that still needs raw scale. The open question is whether that slice shrinks as orchestration matures, or whether multi-step exploit chains and long-horizon reasoning stay a frontier-only zone regardless of how good the surrounding tooling gets.

Does the small-model-plus-orchestration edge hold as the task gets harder, or does multi-step exploitation pull back toward frontier scale?
None of these seven questions has a settled answer — that's simply where the field stands this far into the exploitation era J.P. Morgan's report describes. The capabilities are real and deployed; the measurements, the economics, and the safeguards are still catching up. Benchmarks for machine-speed vulnerability discovery lag the thing they're meant to measure, the maintainer side of the funnel has no tested remedy, and the line between a fabricated exploit and a confirmed one is being drawn by tools that are themselves part of the same automation wave. What the rest of this post argues is that this is exactly why prioritization, not volume, deserves the attention — the open questions are about how fast and how broadly the capability lands, not whether it lands at all.

09 Conclusions

J.P. Morgan's Patchmageddon does the industry a service by putting numbers behind a shift that's been visible anecdotally for a year: AI-driven vulnerability discovery is real, the collapse in time-to-exploit is real, and the old ninety-day disclosure cadence describes a world that no longer exists. What hasn't changed is that prioritization, not raw speed, is what determines whether any of this ends up mattering to a given organization.

Panic is not a strategy.

The empirical reality — Perez and Livshits on uneven exploitation rates, Root Evidence on breach data, Parla's direct test of EPSS against confirmed exploits — is that risk stays highly concentrated even as raw volume explodes.

Over 98% of disclosed vulnerabilities never lead to a breach, and the attackers who do get in keep relying on the same familiar, unpatched, exposed entry points they always have.

The pragmatic summary:

  1. Acknowledge the discovery. Machine-speed vulnerability discovery is real — J.P. Morgan is right about this.
  2. Reject volume panic. "Patch everything, faster" is not a strategy.
  3. Anchor in reachability. Prioritize on reachability and KEV data, not CVSS or EPSS scores alone.
  4. Execute what's reachable. Point rapid remediation at the high-context exposures that are actually reachable.
  5. Layer in prediction. Where nothing has been confirmed exploited yet, a calibrated forward-looking model — Suciu et al.'s functional-exploit predictor, not CVSS and not EPSS alone — adds real signal on top of reachability and KEV data, not in place of them.

The organizations that come out ahead in the exploitation era won't be the ones trying to patch all 48,000 CVEs disclosed this year. They'll be the ones that adopt evidence-based prioritization, measure dynamic reachability, and point their rapid-response capacity at the vital minority of flaws — something under 2% of the total — that actually drive enterprise risk. Everything else is potentially expensive noise, however fast the machine generating it has gotten.

References

The J.P. Morgan Report
On the Collapsing Exploit Timeline
On Frontier Model Capability
On Reachability and Exploitation Prediction
On the Q1 2026 Breach Data
On Cyber Insurance
On the Broader Ecosystem
On Prior Exploitation-Prediction Research
Open Research Questions