Menu
AI & EMERGING TECH

Agentic AI Red Team Checklist Adds 222 Security Tests

Uday Patil Oct 11, 2026 12 min read 11 views
Agentic AI Red Team Checklist Adds 222 Security Tests

Agentic AI red team checklist research is expanding how security teams test autonomous AI systems, with a new free framework providing 222 security tests across 20 attack categories.

The checklist was created by security researcher Ravi Rajput and follows a structure inspired by the OWASP Web Security Testing Guide.

Instead of focusing primarily on prompt injection, the framework pushes testers to examine the entire agentic stack, including infrastructure, cloud identity, model supply chains, tool permissions, memory systems, agent-to-agent communication, Model Context Protocol servers and deployment pipelines.

Each test includes an objective, testing procedure, suggested tools, expected result, severity rating and evidence method so red teams can document not only what they tested, but how they proved the result.

For broader coverage of autonomous-agent risks and defensive controls, see our AI-Era Threats and Agentic Security Guide.

Key takeaway: Agentic AI security testing cannot stop at manipulating prompts. The most serious weaknesses may exist in cloud roles, vector databases, agent tools, MCP servers, orchestration systems and other trust boundaries surrounding the model.

What Is the Agentic AI Red Team Checklist?

The checklist is a spreadsheet-based security assessment framework containing 222 individual tests designed for authorized agentic AI penetration testing and internal red-team engagements.

Rajput describes the project as a response to a common failure in AI security assessments: teams spend significant time testing prompt injection while overlooking traditional infrastructure or access-control weaknesses that can have far greater impact.

Examples include:

  • An exposed MLflow server
  • A reachable cloud metadata endpoint
  • Over-permissioned IAM roles
  • Missing tenant isolation in a vector database
  • Unsafe tool combinations
  • Unauthenticated MCP servers
  • Weak agent-to-agent trust boundaries

The checklist is designed to force testers to examine those areas systematically rather than relying on ad-hoc testing.

222 Tests Cover 20 Attack Categories

The framework organizes 222 checks across 20 categories and four broad phases.

The early stages focus on understanding the environment and identifying what the agent can access before testers begin more aggressive exploitation.

Later stages examine areas such as:

  • Infrastructure and orchestration
  • Cloud identity and permissions
  • Model and software supply chains
  • Prompt injection
  • Unsafe output handling
  • Tool execution
  • Excessive agency
  • Memory and RAG systems
  • Agent networks
  • MCP servers
  • CI/CD pipelines
  • Privilege escalation
  • Lateral movement
  • Persistence
  • Data exfiltration
  • Resource exhaustion
  • Integrity attacks
  • Voice and multimodal inputs

The purpose is to test the system as an interconnected application rather than treating the language model as the only security boundary.

Why Testing Order Matters

One of the checklist’s core ideas is that agentic AI red teaming should follow a deliberate order.

Rajput recommends mapping the environment before attempting prompt injection or other model-level attacks.

This matters because the same prompt-injection weakness can have dramatically different consequences depending on the agent’s permissions.

For example, an injected instruction that only changes generated text may have limited impact.

The same instruction becomes much more serious if the agent can call a tool connected to an administrator-level cloud role.

Security teams therefore need to understand the agent’s infrastructure, identity and accessible tools before accurately rating the risk of model manipulation.

75 Tests Are Rated Critical

The checklist assigns proposed severity levels to all 222 tests.

The current distribution is:

  • 75 Critical
  • 108 High
  • 30 Medium
  • 9 Low

These ratings describe the potential severity of each test scenario.

They do not mean that every agentic AI platform contains 75 critical vulnerabilities.

A finding exists only if the tested system actually fails the relevant check.

Critical Risks Often Exist Outside the Model

Several of the checklist’s highest-impact scenarios involve traditional infrastructure or application-security failures rather than LLM behavior.

Examples highlighted by Rajput include:

  • Cloud credential theft through metadata-service SSRF
  • Unsafe Python pickle deserialization leading to code execution
  • Tool composition that converts permitted reads into unauthorized exfiltration
  • Cross-tenant vector database access
  • Forged agent-to-agent messages
  • Confused-deputy privilege abuse

These weaknesses occur at trust boundaries surrounding the model.

This is one reason agentic AI security testing increasingly resembles a combination of web application testing, cloud security, identity testing and AI-specific adversarial evaluation.

Cloud Metadata Access Gets Special Attention

Cloud metadata endpoints are an important test area because autonomous agents may execute inside workloads that have IAM roles or other cloud identities attached.

If an attacker can manipulate an agent into contacting a metadata endpoint, the resulting credentials may provide access to real cloud resources.

This type of issue resembles the broader risk demonstrated in recent AI-agent research where prompt manipulation combined with broad cloud permissions dramatically increased the blast radius.

For additional context, see our AgentCorruption AWS AgentCore security analysis.

Memory and RAG Systems Need Their Own Security Tests

The checklist treats memory and retrieval systems as independent security boundaries.

One representative test examines whether removing or bypassing a tenant_id filter allows a user to retrieve documents belonging to another customer.

That type of failure is fundamentally different from a hallucination or unsafe model response.

It represents a direct isolation failure between customers.

Agentic systems using vector databases should therefore test:

  • Tenant isolation
  • Authorization filters
  • Retrieval scope
  • Memory poisoning
  • Cross-session contamination
  • Unauthorized document access

MCP Servers Are Part of the Attack Surface

Model Context Protocol servers are another major testing area.

MCP can connect AI agents to tools, files, APIs and enterprise applications.

That flexibility also expands the trust boundary.

Security teams need to assess whether MCP servers:

  • Authenticate clients correctly
  • Restrict sensitive tools
  • Validate arguments
  • Prevent unauthorized actions
  • Protect credentials
  • Resist malicious or manipulated tool descriptions

OWASP’s agentic security guidance similarly emphasizes least-privilege tool access, trusted instruction boundaries and human confirmation before irreversible actions.

Agent-to-Agent Communication Creates New Risks

Multi-agent architectures introduce another class of security problems.

An agent may trust instructions or identity claims received from another agent without sufficiently verifying where they originated.

The checklist includes scenarios involving forged inter-agent messages and delegated privilege abuse.

An attacker who compromises a lower-privilege agent may attempt to convince a more powerful component to perform an action on its behalf.

This resembles the traditional confused-deputy problem, but it now occurs across autonomous components exchanging natural-language instructions and tool requests.

AI Agents Need Supply-Chain Testing Too

Agentic systems can depend on far more than application code.

Their effective supply chain may include:

  • Foundation models
  • Model weights
  • Python packages
  • Plugins
  • MCP servers
  • Prompt templates
  • Tool descriptions
  • Agent configuration

Microsoft’s updated agentic AI failure taxonomy specifically recommends inventorying those components and treating tool descriptions and other natural-language control surfaces as part of the security supply chain.

Prompt Injection Is Still Important, But It Is Only One Category

The checklist does not minimize prompt injection.

Instead, it places prompt injection inside a much larger system-security assessment.

Security teams still need to test direct and indirect prompt injection, instruction conflicts and data-driven manipulation.

But those tests should be combined with questions such as:

  • What tools can the agent invoke?
  • Which cloud permissions does it hold?
  • Can it reach internal services?
  • Can it modify persistent memory?
  • Can it delegate tasks to other agents?
  • Can it access another customer’s data?

The answers determine whether successful prompt injection is merely annoying or becomes a serious enterprise compromise.

Each Test Includes an Evidence Method

A notable feature of the checklist is its evidence classification system.

Tests can produce three broad types of evidence:

  • Reflective: the result appears directly in the agent’s response.
  • Blind: the tester infers success from timing, behavior or a state change.
  • Out-of-band: the target contacts a controlled DNS, HTTP or similar callback service.

The checklist recommends classifying a finding as Confirmed only when reflective or out-of-band evidence is available.

A result supported only by blind evidence should remain Probable.

This helps prevent red teams from overstating findings that cannot be directly proven.

Why Evidence Discipline Matters

AI security findings can be particularly difficult to reproduce because model behavior is probabilistic.

A prompt may succeed once and fail later.

By attaching screenshots, logs, callback evidence or other artifacts to each result, testers can create a stronger audit trail.

This makes the report more useful to engineering teams that need to reproduce and remediate the issue.

The Checklist Maps to Multiple Security Frameworks

Each test includes mappings to established AI-security frameworks.

Rajput says the spreadsheet includes references to:

  • OWASP Agentic Security Initiative
  • OWASP Top 10 for LLM Applications
  • OWASP MCP Top 10
  • OWASP machine-learning supply-chain guidance
  • MITRE ATLAS

This allows organizations to filter or organize test results around the framework used in their internal security program.

However, the checklist author warns that framework identifiers can change over time and should be verified before being used in formal reporting.

The Checklist Is Not an Official OWASP Standard

This distinction is important.

The checklist is independently created by Ravi Rajput.

It uses a structure inspired by OWASP’s Web Security Testing Guide and maps tests to OWASP resources, but it is not an official OWASP checklist or certification standard.

Organizations should therefore treat it as practical community red-team tooling rather than an official compliance requirement.

OWASP Also Calls for Lifecycle-Wide Agentic Red Teaming

Although this particular checklist is independent, its broader philosophy is consistent with OWASP’s current agentic AI red-teaming guidance.

OWASP says autonomous AI systems introduce risks including prompt injection, agent privilege escalation, data poisoning, model misuse and emergent behavior that require coordinated testing across the AI lifecycle.

This reinforces the idea that AI security assessments need to move beyond isolated prompt tests.

Microsoft’s Failure Taxonomy Shows the Same Shift

Microsoft’s 2026 update to its taxonomy of agentic AI failure modes similarly focuses on system-level attack chains rather than model behavior alone.

The company recommends security teams test areas including:

  • Agent supply chains
  • Identity verification
  • Human approval bypass
  • Session-context contamination
  • Inter-agent trust escalation
  • Capability disclosure
  • Goal hijacking

Microsoft specifically notes that some of these attacks cannot be discovered through model-level evaluation alone and require full task-flow testing.

Written Authorization Is Mandatory

The checklist contains tests capable of accessing credentials, executing code, reading cross-tenant data and interacting with cloud infrastructure.

Those activities can be destructive or illegal outside an authorized assessment.

Rajput explicitly states that the checklist is intended only for systems the tester owns or has written authorization to assess.

Potentially destructive tests should be defined in the scope agreement and may need to be performed only in staging environments.

Testers Should Mark Out-of-Scope Checks Explicitly

The checklist recommends setting excluded tests to N/A before the engagement begins.

This creates a record showing that the check was intentionally excluded rather than accidentally forgotten.

Available status values include:

  • Not Started
  • In Progress
  • Passed
  • Failed
  • Blocked
  • N/A

The evidence column can then contain references to screenshots, proxy logs, callback logs or test values.

CVE Numbers Are Deliberately Excluded

The first version of the checklist avoids hard-coded CVE numbers.

The author argues that attack classes remain useful longer than individual vulnerability identifiers.

Security teams are advised to verify current CVE and NVD information separately before including specific identifiers in a final report.

What Security Teams Should Test First

For organizations deploying autonomous AI agents, the practical lesson is to start with architecture and permissions before attempting sophisticated prompt attacks.

Security teams should first identify:

  1. Every deployed AI agent.
  2. The identity assigned to each agent.
  3. The tools each agent can invoke.
  4. External and internal services it can reach.
  5. Memory and vector stores it can access.
  6. MCP servers connected to the agent.
  7. Other agents it can communicate with.
  8. Cloud and CI/CD permissions available to the runtime.

That inventory determines the real impact of later adversarial testing.

Frequently Asked Questions

What is the Agentic AI red team checklist?

It is an independent spreadsheet-based security testing framework created by Ravi Rajput containing 222 test cases across 20 agentic AI security categories.

Is it an official OWASP checklist?

No. It is structured similarly to the OWASP Web Security Testing Guide and maps tests to OWASP frameworks, but it is independently developed.

Does it only test prompt injection?

No. Prompt injection is only one part of the framework. The checklist also covers cloud infrastructure, tools, memory, MCP, agent communication, CI/CD, privilege escalation, exfiltration and other risks.

How many critical tests are included?

The checklist currently rates 75 tests Critical, 108 High, 30 Medium and nine Low.

What does Reflective evidence mean?

Reflective evidence means the successful result is visible directly in the agent’s output or response.

What is out-of-band evidence?

Out-of-band evidence is generated when the tested system contacts an external callback endpoint controlled by the authorized tester.

Can the checklist be used against any AI service?

No. It should only be used against systems the tester owns or has explicit written authorization to assess.

Why is testing cloud identity important?

An AI agent with excessive IAM permissions can turn a prompt or tool-manipulation weakness into credential theft, data access or broader cloud compromise.

Final Takeaway

The new Agentic AI red team checklist reflects an important shift in AI security testing.

Autonomous agents are no longer isolated language models. They operate with cloud identities, tools, memory, APIs, MCP servers, deployment pipelines and connections to other agents.

A security assessment that only asks whether the model can be prompt-injected can therefore miss the vulnerabilities with the greatest real-world impact.

By organizing 222 tests across infrastructure, model, tool and agent trust boundaries, the checklist provides red teams with a structured way to investigate the full system and document exactly what they were able to prove.

Stay Updated on AI Agent Security

Agentic AI is rapidly expanding the traditional enterprise attack surface.

Follow CyberUpdates365 for verified AI security research, agentic threats, MCP vulnerabilities, cloud-security issues and practical defensive guidance.

Primary and Supporting Sources

Ravi Rajput — Original Checklist:
An Agentic AI Red Team Checklist: 222 Tests, WSTG-Style

OWASP GenAI Security Project:
AI Security Solutions Landscape for AI and Agentic Red Teaming

Microsoft AI Red Team:
Taxonomy of Failure Modes in Agentic AI Systems

Uday Patil
About The Author

Uday Patil

Uday Patil is a Cybersecurity Researcher, DevSecOps Engineer, and the Founder of CyberUpdates365. Specializing in Threat Intelligence and Zero-Day vulnerability analysis, Uday is dedicated to breaking down complex cyber threats into actionable insights. His mission is to empower developers, security teams, and aspiring tech talent with rapid alerts, practical guidance, and career mentorship.