CAAP
A repeatable security-testing layer for autonomous AI agents.
CAAP expands the Cogensec Agent Attack Patterns standard into a working 200-pattern taxonomy and a safe, framework-neutral benchmark runner. Two hundred patterns across eleven domains — each with stable IDs, definitions, mappings, and a safe path to an executable test.
A shared vocabulary for how agents fail.
CAAP decomposes broad agentic risk categories into identifiable attack patterns that can be implemented as safe evaluations, regression tests, red-team scenarios, and assurance evidence.
Each pattern carries attack families, a stable ID, a definition, maturity, informative mappings, relationships, severity, and metadata — canonical JSON plus generated YAML and a human-readable page.
The benchmark runner is deliberately framework-neutral. Adapters translate a test case into a target invocation and translate the trace back into normalized events, keeping the taxonomy and oracle model stable across any agent framework.
The public suite uses authorized targets, synthetic data, mock tools and sinks, harmless sentinels, and ephemeral state. It contains no destructive payloads, real exfiltration, production persistence, or approval-bypass procedures.
Eleven domains of agentic attack.
Every pattern maps to a primary OWASP Agentic Security Initiative category.
Goal & Instruction Hijacking
Objective override and instruction injection across trust boundaries.
Tool Misuse & Exploitation
Descriptor poisoning, unsafe tool selection, and argument abuse.
Identity & Privilege Abuse
Impersonation, token scope, and privilege escalation.
Agentic Supply-Chain Attacks
Compromised tools, models, and dependencies in the agent chain.
Unexpected Code Execution
Unintended code paths and sandbox escapes.
Memory, RAG & Context Poisoning
Persistent memory injection and retrieval poisoning.
Insecure Inter-Agent Communication
Spoofed or tampered agent-to-agent messages.
Cascading & Systemic Failures
Fault propagation across coupled agent systems.
Human-Agent Trust Exploitation
Manipulating human oversight and required approvals.
Rogue & Emergent Agent Behavior
Unintended autonomous and emergent action.
Embodied & Physical-Agent Attacks
Physical actuation and sensor-channel manipulation.
One mechanism. One stable identity.
Stable IDs follow CAAP-{DOMAIN}-{NUMBER}. Existing IDs and titles are immutable; a materially different adversarial mechanism receives a new ID. A new prompt, carrier, model, framework, or wording alone does not.
Test variants never receive new pattern IDs — they use a case identifier such as CAAP-GH-02-REF-001.
Direct Objective Override
Tests whether direct objective override can cross an agent trust boundary and cause unauthorized behavior in the direct goal manipulation attack family.
Specification confidence and implementation are separate.
Maturity describes how settled a record is. Implementation status describes whether an enabled public test exists. The two advance independently.
Reviewed public pattern with an enabled, executable reference case. Preserved from CAAP v1.0.0-draft.1.
A designed record seeking implementation and working-group review before it becomes executable.
A stable working ID and definition; implementation is scaffolded and awaiting a contributor case.
From declarative case to five-state result.
A validated case crosses the adapter boundary to an authorized target. Consequential tools are replaced with mocks; normalized telemetry feeds the oracles, which resolve exactly one result state.
Taxonomy and case files are declarative and validated before any execution.
The adapter is the only component that communicates with a target.
Mock tools and sinks form the single permitted side-effect boundary.
Reports retain only normalized evidence supplied by the adapter.
Every run resolves to exactly one state.
All required secure-behavior oracles satisfied and no attack-success oracle fired.
At least one attack-success oracle is satisfied.
Required evidence is absent, contradictory, or insufficient. Never a pass.
The target or harness did not execute the intended scenario.
The target lacks a required capability or trust boundary.
Inconclusive, test error, and not applicable never silently improve it.
Always published alongside the score and result counts.
Baseline severity as the weight across pass and fail.
Reachability, proven with harmless sentinels.
CAAP demonstrates that a weakness is reachable — it does not ship destructive payloads. The guardrails are built into the harness, not left to the operator.
The vulnerable mock exits nonzero as each synthetic oracle fires. No real side effect occurs — the sink accepts only a CAAP sentinel token and stores it in memory.
Toward a stable community benchmark.
Dates are intentionally omitted; safety and independent validation determine readiness.
Public review foundation
Independent adapters
Profiles & composition
Stable community benchmark
Test your agents against the patterns that matter.
Get in touch to run CAAP against your agents. We’ll scope an authorized assessment, run the suite, and walk your team through the evidence.

