Palo Alto Turned Mythos Into a Pen Test That Never Stops
Unit 42's new subscription service runs Anthropic's Claude Mythos, OpenAI's GPT-5.6-Cyber and open-weight models against customer networks around the clock, on the theory that no single model sees enough on its own.
The most restricted AI models in cybersecurity now come with a sales team. On September 22, Palo Alto Networks announced Unit 42 Continuous Frontier AI Defense, an offensive security service that runs Anthropic's Claude Mythos and OpenAI's GPT-5.6-Cyber against a customer's systems on a continuous basis, finding, validating and helping fix exposures before an attacker gets to them. Palo Alto itself describes both as gated capability models, not tools anyone can sign up for. Palo Alto's pitch is that enterprises can get the capability through a vendor they already pay, with Unit 42's red teamers steering it.
The product is a subscription, not a one-off engagement. That distinction is the whole story. Penetration testing has traditionally been an annual or quarterly event: a team arrives, spends a few weeks attacking, writes a report and leaves. Palo Alto is betting that model-driven attackers have made that cadence obsolete.
What Palo Alto is actually selling
Per the company's press release, the service starts with a baseline scan of a customer's full estate, then keeps testing. It covers web applications, APIs, cloud infrastructure, source code and networks. Findings come with prioritized fixes, code-level guidance and virtual patch recommendations, so a team can block an exploit path while a real fix is written.
The core engineering piece is what Palo Alto calls a proprietary multi-model harness. In a launch post, Unit 42 SVP Sam Rubin describes it as an orchestration layer that routes each offensive testing task to the model best suited for it. The model roster is Claude Mythos 5, GPT-5.6-Cyber and unnamed open-weight models. Pricing is annual and varies depending on which OpenAI, Anthropic and open-source models a customer chooses. Palo Alto did not publish figures.
The company says it spent $17 million over six months on research and methodology for the service and validated it across more than 100 Unit 42 customer engagements, per the release and TelecomTV's report.
The numbers, and whose numbers they are
Palo Alto's headline claim is that continuous Mythos-based scanning, in internal deployment, produced a year's worth of traditional penetration testing results in three weeks and found 3.2 times more high and critical vulnerabilities per product than legacy methods. In customer assessments, the company says it found exposures in 100% of tested organizations, rated 37% of those exposures high or critical, and that more than 66% of third-party application exposures had no known CVE.
These are vendor figures from a launch announcement, and nobody outside Palo Alto has audited them. The last one is still the most useful. An exposure with no CVE is one that a patch-management program, which works from lists of known vulnerabilities, will never flag. If two in three of the problems Palo Alto finds in third-party apps fall into that category, a lot of real risk sits outside the process most companies use to measure their security.
Rubin's post includes a customer example from financial services. The model-driven testing chained three separate weaknesses: identity re-verification failing on payment links, skipped one-time-password checks, and a session routing flaw. Together they allowed complete account takeover and payment fraud with no action required from the victim. None of the three is dramatic alone. Chaining small flaws into a working attack is the kind of work that used to require a skilled human tester with time to spare.
Why one model is not enough
The most interesting claim in the launch is an admission. Per Rubin's post, no single AI model catches more than 40% of vulnerabilities in complex environments, and leading cyber models have less than 10% overlap in the exposures they identify.
If that holds up, it matters more than the marketing numbers. It means the frontier cyber models are not interchangeable versions of the same capability. Each finds a largely different set of problems. The practical conclusion is that a defender relying on one model has most of its attack surface unexamined, and that the value sits in the harness that combines them. That is also, conveniently, the part Palo Alto owns. Both things can be true.
It also explains why the labs are willing to partner. McCall McIntyre, OpenAI's head of global cyber partnerships, said in the release that defenders need frontier capabilities "within the tools, workflows, and services they already trust." Anthropic's cybersecurity lead, Michael Moore, said Claude Mythos "found flaws that survived decades of human review," and more than ten thousand high-severity vulnerabilities across widely used software. For labs that have restricted these models because of their offensive potential, a large security vendor with existing customer contracts and incident-response staff is a controlled distribution channel. That reading is our inference; neither company framed it that way.
The speed argument
Palo Alto's justification is time. The company says threat actors using AI have compressed breach cycles by nearly 97% in some cases, from weeks to hours. Rubin's post cites an intrusion completed using more than 50 MITRE ATT&CK techniques in under 10 hours. "AI has created an asymmetric advantage for threat actors against organizations trying to defend at human speed," Rubin said in the release.
The logic follows: if an attacker can find and chain flaws in hours, a test that runs once a year describes a network that no longer exists. Continuous testing is the obvious response.
The obvious risk is also there. A service that constantly runs the most capable offensive models against production systems is itself a sensitive asset, and it produces a live map of a company's weaknesses. Palo Alto stresses that the work is led by its offensive security experts, not left to run unsupervised. Buyers should ask how that supervision works in practice, where findings are stored, and what the models are allowed to touch.
What to watch
First, whether competitors follow with their own multi-model offerings, which would confirm Palo Alto's claim that coverage comes from combining models rather than picking the best one. Second, whether Anthropic and OpenAI open similar arrangements to other security vendors, or keep gated cyber models in a few hands. Third, whether any customer publishes independent results. Until then, the 40% ceiling and the under-10% overlap are the figures worth remembering, because they describe the limits of the tools, not just their strengths.
