Key Takeaways
- AI agents are becoming potent offensive tools, creating a new, constantly expanding attack surface that traditional defenses struggle to monitor.
- Treating every AI agent as a privileged identity and adopting zero‑trust, phishing‑resistant authentication are essential first steps for defenders.
- Continuous, AI‑native red‑team and penetration‑testing capabilities are emerging as a critical countermeasure to match the speed and scale of AI‑driven attacks.
- Real‑world demonstrations (e.g., Armadin’s “largest controlled live AI cyberattack”) show that autonomous agent swarms can uncover millions of offensive actions and dozens of zero‑day‑class vulnerabilities in hours rather than months.
- Human security professionals must evolve from routine testers to experts who assess business impact, understand AI behavior, and design guardrails for increasingly AI‑centric infrastructures.
The Rising Threat of Agentic AI
AI agents have moved beyond generating content to executing actions inside enterprise environments, and they have already been observed in real‑world attacks. As Matt Hartman, former acting head of cyber at CISA, warned, every agent must be treated as a privileged identity because it can gain access to sensitive systems and data. This shift introduces a fresh attack surface that blends human and machine intrusions, giving defenders yet another vector to monitor.
New Attack Surfaces and Identity Challenges
Beyond direct compromise, AI agents create novel data‑integration channels that adversaries can abuse. They also spawn a rapidly growing class of non‑human identities that are difficult to inventory and often slip past static security policies. Hartman emphasized that organizations need strong identity controls, phishing‑resistant authentication, behavioral analytics, and a zero‑trust mindset to cope with these changes—though the fundamentals are not new, the scale is.
Attacker Advantages: Tireless, Focused Bots
From the attacker’s perspective, AI agents never tire, take holidays, or lose focus. They can autonomously hunt for vulnerabilities, map networks, and locate sensitive files at machine speed, making them invaluable assets for financially motivated criminals and state‑sponsored teams. This relentless capability has turned agentic bots into a “gift from the heavens” for offensive operators.
Agentic Red Teaming as a Defensive Countermeasure
Recognizing the asymmetry, experts like former NSA cyber chief Rob Joyce advocate turning the same technology inward: if you aren’t using AI agents to attack your own systems, someone else will. Hartman echoed this, noting a burgeoning market for continuous, AI‑native, automated red‑team and penetration‑testing tools that let organizations stay ahead of threats by constantly probing their defenses.
From Government to Private Sector: Hartman’s Role
After nearly two decades at CISA, Hartman joined Merlin Group as chief strategy officer. In this position he helps identify early‑ to growth‑stage cybersecurity firms worthy of investment and guides them through government, critical‑infrastructure, and other regulated markets. His current focus is on integrating promising AI‑driven technologies—especially those that pit AI agents against each other—into high‑assurance environments.
Armadin’s Massive Live AI Cyber‑Attack Demonstration
Mandiant founder Kevin Mandia’s new venture, Armadin, launched in March with $190 million in seed and Series A funding. The company builds autonomous attacker swarms—thousands of AI agents that run 24/7 inside client infrastructure to emulate real attackers. Together with Tenex.ai, Armadin conducted what it called the “largest controlled live AI cyberattack on record” for an unnamed global institution. Over three days, the swarm generated 17 million offensive actions, uncovered 38 validated attack paths, and produced 238 security findings. Tenex.ai’s platform correlated 101,169 alerts across 231 billion raw events, a task that would have required a five‑person analyst team roughly 2,400 hours (about four months) to complete manually.
Why Autonomous Agents Outperform Human Teams
Evan Peña, Armadin’s co‑founder and chief offensive security officer, previously led Mandiant’s 210‑person global red‑team. He noted that traditional, human‑led assessments are limited by time and workforce constraints—typically a few weeks or a month per engagement, followed by a lengthy hiatus before retesting. AI agents eliminate those limits: they never sleep, need no holidays, and can be pre‑ and post‑trained with deep expertise in coding, source‑code review, application security, network misconfigurations, and exploitation techniques. This combination yields vastly greater coverage; where a human team might assess only 1,000‑2,000 of 10,000 external systems in a given period, an agent swarm can scan all 10,000 in a matter of hours.
Real‑World Impact: Zero‑Days Found at Scale
Peña claimed that Armadin’s agents have broken into every customer environment they have tested, uncovering over 50 zero‑day vulnerabilities. Importantly, he distinguished high‑impact zero‑days—those granting remote code execution on actual production systems—from low‑severity flaws that merely deface a webpage. The focus is on impact, not noise, underscoring the potency of AI‑driven offensive capabilities.
The Inadequacy of Periodic Pen‑Testing
Jay Bavisi, founder and group president of EC‑Council, argued that legacy penetration‑testing cadences—annual for compliance, quarterly for the more diligent—are obsolete in the face of AI‑accelerated threats. Human‑led tests often take three months, leaving a dangerous speed gap, and they rarely cover the entire organization. Attackers, by contrast, employ AI to continuously scan every asset, bypassing scope and sophistication limitations that plague manual teams.
Upskilling for the AI Era
To bridge the gap, EC‑Council has launched a sponsored CPENT AI examination for pen‑testing professionals, pairing successful candidates with $1,000 in cybersecurity training credits for nonprofit partners. Bavisi stressed that the traditional pen‑tester role will not disappear; instead, it will evolve into a broader discipline that includes assessing business impact, prioritizing remediation, and mastering the testing of large language models (LLMs) and agentic behaviors. Professionals must now understand harm taxonomies, evaluate guardrails, and make engineering decisions about what constitutes acceptable risk in AI‑integrated systems.
The Future Role of Security Professionals
As AI becomes the “heartbeat” of organizations, offensive AI security specialists will be tasked with verifying the robustness of those systems. Pen‑testers will need to think like attackers, anticipate how agents might misuse AI capabilities, and design resilient defenses that go beyond signature‑based detection. In short, the job has expanded from periodic vulnerability checks to continuous, impact‑focused validation of an AI‑laden attack surface.

