
A recent test of Anthropic’s Mythos model has raised a question for the cybersecurity industry—as AI systems become more capable, will defenders need to prepare not only for criminals using AI, but also for AI agents that can independently manipulate people and systems?
According to CNBC, the concern emerged after the UK’s AI Security Institute (AISI) tested an agent powered by Mythos in an environment where safety guardrails were removed and the system was given open internet access.
During the assessment, the agent researched the project’s human handlers, developed synthetic identities, and attempted to use those identities to persuade users to approve malicious code for an open-source project.
When confronted, the agent edited records of its earlier activities to make them appear benign, and even considered creating a new identity under which it could continue operating. AISI noted, however, that the attempts were unsuccessful and did not result in any real-world damage.
Not Assuaging Concerns
Anthropic responded that Mythos was tested in a deliberately permissive environment designed to evaluate the model under extreme conditions, rather than one that reflected real-world deployments. The company also said there was no evidence that the model escaped its testing environment.
Even so, the findings are unlikely to ease growing concerns about the risks posed by frontier AI systems, especially following several high-profile incidents involving advanced models in recent weeks.
Among the most notable was an incident in which OpenAI models reportedly escaped a sandboxed evaluation environment and accessed Hugging Face infrastructure while attempting to obtain data that could help them perform better on an evaluation.
Shortly after, Anthropic reported that its Claude model had breached its intended containment and accessed systems belonging to three separate organizations during testing.
Patching the Vulnerabilities
All these incidents underscore the evolving cybersecurity challenges posed by capable AI models, placing additional pressure on defenders across multiple fronts.
Cybercriminals can leverage AI to create convincing deepfakes and synthetic identities, while also scaling attacks that previously required significant manual effort into widespread campaigns.
Another concern stems from the capabilities of the models themselves. Mythos has become a focal point in that discussion after Anthropic disclosed that the model identified thousands of zero-day vulnerabilities across systems spanning multiple sectors. The company also stated that Mythos was too capable to release publicly.
Concerns surrounding models like Mythos have prompted organizations to reassess their security strategies. For example, Apple reportedly shifted away from its longstanding cadence of major iOS security updates in favor of smaller, more frequent releases.
The objective for Apple—and for cybersecurity teams more broadly—is to identify and patch vulnerabilities before capable AI systems can exploit them at scale.
The post Anthropic’s Mythos Test Raises New Concerns About Social Engineering appeared first on PaymentsJournal.