A month ago, Anthropic built their most powerful model ever, then decided not to release it. What they are doing instead, is that they provide controlled access to a handful of big players like Amazon, Google, Apple, Crowdstrike, and the Linux Foundation, through an initiative called the Project Glasswing [1].
The model operates at the level of a senior red-team engineer. It can run reconnaissance, identify vulnerabilities, chain them into working exploits, and adapt on the fly. In testing, it's already uncovered bugs that had gone unnoticed by humans for years.
As with every powerful technology we've seen in history so far, the same tool can be used to secure the internet, or it can be used as a weapon. Anthropic is well aware of that, and they describe Mythos as their best-aligned model so far but also as their greatest alignment risk. Their intention is that Mythos is used defensively, so as to harden critical infrastructure before equivalent capabilities show up in open source models. That window is 1-2 years, maybe less.
Making advanced technology accessible to anyone sounds like the way forward if the goal is real democratization. Of course, democratization here could quickly turn into a nightmare if the tool falls into the wrong hands, and it would. Nevertheless, we can't ignore that feeling of discomfort when only big players have early access to such powerful technology and get their code hardened, while the rest — academia, smaller orgs, open-source projects — just fall behind.
Another worrying thought is what happens when (not if) other orgs, potentially guided by very different moral compasses than Anthropic, develop equally powerful models. Worth to mention, Claude Security [2] (being powered by Opus 4.7 [3] rather than Mythos) is Anthropic's way of getting defenders tooled up before that happens.
Finally, speaking of alignment, Mythos might be the best-aligned model ever, and Anthropic has made a strong effort on interpretability with this model, but doesn't its remarkable agentic, creative, and thorough behavior inherently carry the potential for catastrophic misaligned actions? In other words, does best effort ensure zero failure? This is a point they make themselves in the model's own system card, using an analogy with a mountaineering guide to explain the paradox of how a more skilled guide can sometimes put followers in greater danger [4].
We see Mythos as a new threat actor class. Debating whether Mythos can replace your red team is the wrong question to ask, and I will explain why.
The attack surface is expected to explode when we are dealing with an attacker that never sleeps or gets bored, and can easily probe thousands of paths in parallel. Maybe then, vulnerabilities that were low-prio because no human could realistically find them will become exploitable overnight.
If your org deploys AI agents, human red-teamers writing prompt injections by hand won't scale. You'd need an automated system that generates and escalates attacks against your agents.
Having access to a powerful model is not the same as having red teaming capability in your org, which is necessary to integrate the tool to your specific stack (a customer service agent with email access has a completely different attack surface than a coding agent with CI/CD access), maintain pipelines, implement orchestration logic, set up evaluation, calibration and judging, and produce coverage analysis, and decide on when/how iterations will be ran. Good tools do not equal capability, we still need people who know how to use powerful tools.
Maybe things will get harder for red teams once tools like Mythos get integrated as features in platforms but if history repeats itself, and it always does, this will provide a generic scanning. Deep quality work stays tied to talent and some good working hours put to it, still.
Good tools need an adoption curve. But in cybersecurity, you don't have the luxury of waiting and taking your time asking whether Mythos can replace your red team. While you wait, agents in production are exposed. The real question is whether you'll do it fast enough.
Mythos is proof that we've crossed from chatty assistants to autonomous agents that can do real damage in the real world. Anthropic deserves credit for not shipping it for hype. To what extent do we really believe that any single lab can gatekeep this level of power forever?
A month ago, Anthropic built their most powerful model ever, then decided not to release it. What they are doing instead, is that they provide controlled access to a handful of big players like Amazon, Google, Apple, Crowdstrike, and the Linux Foundation, through an initiative called the Project Glasswing [1].
The model operates at the level of a senior red-team engineer. It can run reconnaissance, identify vulnerabilities, chain them into working exploits, and adapt on the fly. In testing, it's already uncovered bugs that had gone unnoticed by humans for years.
As with every powerful technology we've seen in history so far, the same tool can be used to secure the internet, or it can be used as a weapon. Anthropic is well aware of that, and they describe Mythos as their best-aligned model so far but also as their greatest alignment risk. Their intention is that Mythos is used defensively, so as to harden critical infrastructure before equivalent capabilities show up in open source models. That window is 1-2 years, maybe less.
Making advanced technology accessible to anyone sounds like the way forward if the goal is real democratization. Of course, democratization here could quickly turn into a nightmare if the tool falls into the wrong hands, and it would. Nevertheless, we can't ignore that feeling of discomfort when only big players have early access to such powerful technology and get their code hardened, while the rest — academia, smaller orgs, open-source projects — just fall behind.
Another worrying thought is what happens when (not if) other orgs, potentially guided by very different moral compasses than Anthropic, develop equally powerful models. Worth to mention, Claude Security [2] (being powered by Opus 4.7 [3] rather than Mythos) is Anthropic's way of getting defenders tooled up before that happens.
Finally, speaking of alignment, Mythos might be the best-aligned model ever, and Anthropic has made a strong effort on interpretability with this model, but doesn't its remarkable agentic, creative, and thorough behavior inherently carry the potential for catastrophic misaligned actions? In other words, does best effort ensure zero failure? This is a point they make themselves in the model's own system card, using an analogy with a mountaineering guide to explain the paradox of how a more skilled guide can sometimes put followers in greater danger [4].
We see Mythos as a new threat actor class. Debating whether Mythos can replace your red team is the wrong question to ask, and I will explain why.
The attack surface is expected to explode when we are dealing with an attacker that never sleeps or gets bored, and can easily probe thousands of paths in parallel. Maybe then, vulnerabilities that were low-prio because no human could realistically find them will become exploitable overnight.
If your org deploys AI agents, human red-teamers writing prompt injections by hand won't scale. You'd need an automated system that generates and escalates attacks against your agents.
Having access to a powerful model is not the same as having red teaming capability in your org, which is necessary to integrate the tool to your specific stack (a customer service agent with email access has a completely different attack surface than a coding agent with CI/CD access), maintain pipelines, implement orchestration logic, set up evaluation, calibration and judging, and produce coverage analysis, and decide on when/how iterations will be ran. Good tools do not equal capability, we still need people who know how to use powerful tools.
Maybe things will get harder for red teams once tools like Mythos get integrated as features in platforms but if history repeats itself, and it always does, this will provide a generic scanning. Deep quality work stays tied to talent and some good working hours put to it, still.
Good tools need an adoption curve. But in cybersecurity, you don't have the luxury of waiting and taking your time asking whether Mythos can replace your red team. While you wait, agents in production are exposed. The real question is whether you'll do it fast enough.
Mythos is proof that we've crossed from chatty assistants to autonomous agents that can do real damage in the real world. Anthropic deserves credit for not shipping it for hype. To what extent do we really believe that any single lab can gatekeep this level of power forever?