A year ago, an "AI agent" was usually one model with a few tools. Today it's a swarm: a triage agent that routes to a knowledge-base agent, that escalates to a finance or admin agent, that calls shared tools over MCP. Frameworks like LangGraph, CrewAI and the OpenAI Agents SDK made this the default shape of production AI.
That shape introduced a new attack surface that didn't exist when there was only one agent — and it's exactly the surface most testing tools don't look at.
What the field tests today
The agentic-security space grew up fast in 2025–26, and the tooling is genuinely good — but almost all of it is node-focused: it tests one agent, one model, one boundary at a time.
- Single-model vulnerability scanners — open-source tools like NVIDIA's garak probe a model for prompt injection, data leakage, toxicity and 100+ other categories. Essential, and entirely about one model's responses.
- Agentic red-team platforms — commercial products (e.g. Splx AI, now part of Zscaler, and others surveyed in the 2026 landscape) pair automated red-teaming with runtime guardrails and compliance reporting. They've added agent coverage — prompt injection, tool misuse, privilege escalation — but framed at an agent's system boundary.
- Framework eval harnesses — LangSmith, CrewAI's and others' test tooling check whether an agent works (functional correctness, task success), not whether it can be walked off-task by an adversary.
- Runtime guardrails / firewalls — Lakera, NeMo Guardrails and the "LLM firewall" category filter inputs and outputs in production. Valuable containment — but filtering, not pre-deployment pen-testing.
The standards reflect the same lens. OWASP's Top 10 for Agentic Applications (December 2025) made goal hijack, tool misuse and memory poisoning first-class risks — a big step, still organised around a single agent's surfaces. And the threat is already real: security researchers in early 2026 documented autonomous agents running prompt injections against other autonomous agents.
The gap is the edge, not the node
A swarm is a graph. Its highest-impact vulnerabilities don't live inside any one agent — they live on the edges: the handoffs where one agent passes context, instructions, or delegated authority to another that doesn't independently re-validate it.
The attacker touches the agent with nothing to steal. The damage happens at the one three hops away that can move money — and nobody tested the path between them.
Concretely: you inject at a public, low-trust triage agent. Its output is handed to a privileged finance agent as trusted context. The front agent has no sensitive tools, so a single-agent scan of it finds nothing interesting. The back agent is only reachable "internally," so it's often not scanned as an external target at all. The exploit lives entirely in the handoff — the place neither single-agent test looked.
Two properties make this invisible to node-by-node testing:
- Entry ≠ blast radius. Where the attack enters and where it lands are different agents. Scoring the entry point understates the risk.
- Trust is assumed across the edge. Downstream agents treat upstream output as already-authorised. A low-trust → high-trust, context-passing handoff is a privilege-escalation path by construction.
How AAPT is different: we test the swarm as a graph
AAPT onboards the whole swarm as one target and tests both the agents and the handoffs between them.
- Model the graph. Each agent is a node with a trust tier; each handoff is an edge with a computed trust Δ and a context-passing flag. You describe it in a manifest, or import it straight from an AgentSwarms / LangGraph / CrewAI export.
- Edge-aware probes. We inject at the upstream agent, follow the handoff, and detect at the downstream agent — the attacks single-agent tools can't express. Today that includes handoff instruction injection (does an instruction smuggled through a handoff get obeyed downstream?) and cross-agent privilege escalation (can low-trust input upstream trigger a privileged tool call downstream?).
- Score where it lands. A cross-agent finding is scored with the downstream agent's blast radius and regulatory exposure — so an exploit that reaches a system-wide, regulated finance agent from a public chatbot scores like the critical it is, not like a harmless chatbot bug.
- One run, whole swarm. A single fan-out scans every agent with the full probe suite and walks every handoff, rolling up to per-agent and cross-agent findings with CVSS-A scores.
And the parts that already set AAPT apart still apply to the swarm: CVSS-A risk scoring, a remediation tracker with a verified-clean re-test loop, audit-ready compliance evidence (OWASP, MITRE ATLAS, NIST AI RMF, ISO 42001, SOC 2), and a fully sovereign / on-prem mode where even the judge model runs inside your boundary.
What testing a swarm actually looks like
- Onboard — import your LangGraph/CrewAI/AgentSwarms export (or write a short manifest). AAPT builds the node/edge graph with trust tiers.
- Run — one click scans each agent and every handoff.
- Review — per-agent findings plus the cross-agent ones, each with a CVSS-A score and the exact edge it traversed (e.g.
triage → finance, Δ +3). - Prove the fix — re-test the edge after you harden the handoff, and attach the signed evidence to your audit.
What we won't claim
Same honesty as always: testing is a pre-deployment control, not a runtime shield. Edge-aware probes surface that a handoff can be abused so you can harden it — re-authenticate the real principal at the privilege boundary, stop treating upstream text as authorised, scope tools to the caller. The thing that blocks a live cross-agent attack is runtime containment and least privilege. Testing and runtime are two jobs; AAPT does the first, and the runtime guardrail products above do the second. They're complementary — but you can't harden a seam you never knew was there.
The takeaway
The industry is converging on a real answer for single agents. The attack, meanwhile, has already moved to the space between them — and the place a swarm is weakest is the handoff nobody tested. If you're shipping multi-agent systems, the question worth asking isn't just "is each agent safe?" It's: can an instruction that enters the safe agent come out the dangerous one?