Effective Red Teaming in the Agentic Era: Amazon Bedrock (Part II)

The AI attack surface spans beyond models; other components are similarly important and must be secured. This second part of the Bedrock Red Teaming blog series provides further practical guidance on means to continuously test Bedrock components.
8.10.2026
Kennedy Torkura
6 Minutes
AI Security, AI Red Teaming, AWS Security, series: bedrock red teaming
Mitwirkende
Kennedy Torkura
Kennedy Torkura
Co-Founder & CTO
Vielen Dank! Ihre Nachricht wurde erfolgreich übermittelt.
Hoppla! Beim Absenden des Formulars ist ein Fehler aufgetreten.

This is the second part of the blog series on red teaming for Amazon Bedrock. Part 1 covered several aspects of the Bedrock attack surface, including Bedrock agents and Knowledge Bases. This second part discussed five additional areas: Bedrock AgentCore, prompt management, custom models, action groups, and guardrails.

‍

Bedrock AgentCore

Bedrock AgentCore is a framework-agnostic, model-agnostic agentic platform offering serverless agent runtime, OAuth-based identity, and managed memory capabilities. AgentCore allows organizations to operate agentic infrastructure at scale without building the orchestration substrate from scratch.

‍

Figure 1: AgentCore four-path compromise

‍

‍

Ironically, AgentCore also introduces several security issues, including persistent agent memory, code interpreter abuse, and browser runtime manipulation. We recently did a deep dive into these security issues and highlighted cross-agent privilege escalation weaknesses. Read that blog post here: AgentCore or AgentSore: Cross-Agent Privilege Escalation in Bedrock AgentCore.

Red Teaming Objectives

  • Identity Vault Compromise: Target OAuth refresh tokens held in the identity vault and the confused-deputy paths that let a compromised agent assume identities beyond its mandate.
  • Memory Poisoning: Plant adversarial content in agent memory to steer the agent's behavior or extract data, including content that persists across sessions and bleeds across agents.
  • Tool Descriptor Injection: Exploit gateway tool sprawl by injecting malicious tool descriptors that the agent ingests during tool discovery.
  • Browser Sandbox Escape: Serve crafted web content to the managed browser runtime to break out of its sandbox while the agent is browsing.
  • Code Interpreter Privilege Escalation: Abuse wildcard InvokeCodeInterpreter permissions to run code in any code interpreter in the account, including one whose Identity and Access Management (IAM) role holds more privileges than the compromised agent's.

Related References

Maps to CSA Agentic AI Red Teaming Guide Section 4.6 (Agent Impact Chain and Blast Radius), Section 4.8 (Agent Memory and Context Manipulation), and Section 4.9 (Agent Orchestration and Multi-Agent Exploitation, especially Section 4.9.4 Confused Deputy), T1098 (Account Manipulation), T1059.009 (Command and Scripting Interpreter: Cloud API), AML.T0053 (AI Agent Tool Invocation).

Prompt Management

Bedrock Prompt Management is a managed library for storing, versioning, and deploying the system instructions that direct an agent or application's behavior at runtime. Prompts can be managed via Infrastructure-as-Code (IaC) and the AWS Cloud Development Kit (CDK) for proper change control, but the AWS Console and direct API paths remain open by default and bypass any review pipeline that has been put in place.

‍

Prompts encode a lot: business logic, persona constraints, internal context, explicit security rules, and examples of allowed and disallowed behavior. Reading them exposes proprietary intellectual property and reveals the exact controls an attacker would need to bypass. Prompts also belong to the threat class that CSA research describes for agent context files such as SKILL.md, CLAUDE.md, and AGENTS.md: natural-language instructions an agent is designed to trust and follow, with the model itself acting as the execution engine for whatever they say. A poisoned prompt needs no software vulnerability to work. Because a single managed prompt can serve many agents and applications, tampering with it becomes a supply chain compromise that reaches everything consuming it, with no retraining or redeployment. The difference from agent context files is where the impact lands: a poisoned context file compromises a developer's workspace, while a managed prompt is fetched at runtime by production agents and applications, putting customer-facing behavior and production data in scope.

‍

Table 1: Agent context files versus Bedrock Prompt Management.

‍

‍

Red Teaming Objectives

  • System Instruction Disclosure: Harvest system instructions from prompt resources, for example through GetPrompt and ListPrompts, to expose how the application is instructed to behave, such as the persona, internal rules, filtering criteria, and permissions.
  • Draft and Version Tampering: Tamper with a prompt's working draft or publish a tampered version, including through UpdatePrompt and CreatePromptVersion, altering downstream agent and application behavior without retraining or redeployment.
  • Prompt Supply Chain Poisoning: Tamper with a shared prompt outside the IaC pipeline, through the Console or API, or through the pipeline itself, so the change propagates to every agent and application that consumes it.
  • Sensitive Data Harvesting: Mine prompt bodies for Personally Identifiable Information (PII), secrets, and internal URLs inadvertently stored as context.
  • Prompt Deletion: Delete prompts or prompt versions that production consumers reference, breaking every agent and application that calls them.

Related References

Maps to CSA Agentic AI Red Teaming Guide Section 4.4 (Agent Goal and Instruction Manipulation, especially Instruction Set Poisoning) and Section 4.11 (Agent Supply Chain and Dependency Attacks), T1213 (Data from Information Repositories), T1565.001 (Data Manipulation: Stored Data Manipulation), OWASP LLM07 (System Prompt Leakage), LLM02 (Sensitive Information Disclosure), LLM03 (Supply Chain).

Custom Model Attacks

In addition to managed foundation models, Bedrock supports model customization through fine-tuning and continued pre-training methods. This allows organizations to customize model behavior by consuming training data from S3 while providing inference through the same uniform API as managed models. 

‍

Custom models embed proprietary data and trained behavior in their weights, which makes them both high-value attack targets and data leakage surfaces. Training datasets require strong access controls. Separately, the weights themselves are vulnerable to model inversion attacks, where an attacker reconstructs training data from model outputs. Also, attackers often use Model distillation attacks to clone model behavior; a real exfiltration pattern. We previously discussed poisoning training data in this LinkedIn post.

‍

Figure 2: A model poisoning attack orchestrated by targeting the training data in an S3 bucket.

‍

‍

Red Teaming Objectives

  • Training Data Exploitation: Read, alter, or poison the training datasets in S3 that customization jobs consume, to shape the behavior of the resulting model.
  • Cross-Account Model Exfiltration: Abuse Model Share via AWS Resource Access Manager, or Model Copy jobs, to move a proprietary custom model into an attacker-controlled account or Region.
  • Supply Chain Compromise: Plant a tampered or backdoored model into the deployment pipeline through unvalidated marketplace or imported model artifacts.
  • Model Distillation: Sample the model at scale through the invocation API to clone its proprietary behavior.
  • Model Inversion: Probe a custom model with crafted queries to reconstruct sensitive records embedded in its weights during training.

Related References

Maps to CSA Agentic AI Red Teaming Guide Section 4.11 (Agent Supply Chain and Dependency Attacks), AML.T0024 (Exfiltration via AI Inference API), AML.T0020 (Poison Training Data), OWASP LLM03 (Supply Chain), LLM04 (Data and Model Poisoning).

‍

Action Group Abuse

Bedrock Agents use Action Groups to interact with AWS services and interact with external, non-AWS systems (tools). Action groups are composed of Lambda functions, including assigned execution roles, and agents leverage these to exert autonomy and decision-making. However, attackers abuse this mechanism to maliciously control agents. Attackers can use this approach to take over agents, manipulate agent reasoning, and hijack tools. The blast radius scales with the reach of the action group's execution role.

‍

Red Teaming Objectives

  • Tool Call Hijacking: Plant indirect prompt injection that drives the agent into invoking tools it was never meant to call.
  • Over-Privileged Tool Abuse: Exploit Lambda execution permissions that exceed the agent's declared action scope, the classic gap where the tool can do more than the agent should.
  • PassRole Exploitation: Target weak iam:PassRole scoping on agent provisioning and Lambda execution roles to bind excessive permissions to the tool layer.
  • Reasoning Loop Poisoning: Feed crafted tool responses back into the agent's reasoning loop to steer its next action.
  • Approval Gate Evasion: Force destructive actions such as delete, transfer, or publish through paths that lack a human-in-the-loop approval gate.

Related References

Maps to CSA Agentic AI Red Teaming Guide Section 4.1 (Agent Authorization and Control Hijacking) and Section 4.9.4 (Confused Deputy Attack), T1098 (Account Manipulation), OWASP LLM06 (Excessive Agency).

Guardrail Evasion and Tampering

Bedrock Guardrails are the runtime safety layer that sits between users and foundation models. They evaluate both input (before the model sees it) and output (before the user sees it) across six configurable policies: content filters, denied topics, word filters, sensitive information filters, contextual grounding checks, and Automated Reasoning checks. Guardrails are the runtime checker between user, agent, and model, which makes them a target in their own right. An attacker who manipulates a guardrail removes the protection it provides, while everything behind it still appears protected. 

‍

Guardrails can be evaded with crafted inputs or manipulated directly: switched off, weakened, or set to detect mode, where they still evaluate and report but never intervene. When this happens to a guardrail protecting an agent, you have what the CSA Agentic AI Red Teaming Guide (Section 4.2) calls a Checker-Out-of-the-Loop condition: the policy is technically "on," but the agent operates as if it were absent. Foundation models' built-in safety catches generic harmful content but not organization-specific risk; e.g., Amazon Nova moderates output across safety, sensitive content, fairness, and security.

 

Figure 3: Using Mitigant Attacks to Maliciously Update Bedrock Guardrails policies.

‍

‍

Red Teaming Objectives

  • Filter Jailbreaking: Craft jailbreak and prompt-injection payloads that slip past the prompt-attack filter and reach the foundation model intact.
  • Multilingual Evasion: Phrase malicious requests in languages or encodings that the denied-topic and word filters fail to catch.
  • Multi-Modal Smuggling: Embed instructions as text inside image inputs to carry an injection through a channel the text filters never inspect.
  • Guardrail Tampering: Switch off, weaken, or flip guardrail policies to detect mode, for example through UpdateGuardrail, or publish a weakened version for consumers pinned to a version, so policy-violating prompts pass while the guardrail stays attached.
  • Grounding Bypass: Craft retrieval context that satisfies the contextual grounding check while still carrying the injection through to the model.

Related References

Maps to CSA Agentic AI Red Teaming Guide Section 4.2 (Checker-Out-of-the-Loop), AML.T0051 (LLM Prompt Injection), AML.T0054 (LLM Jailbreak), T1562.001 (Impair Defenses: Disable or Modify Tools), OWASP LLM01 (Prompt Injection).

Secure Your AI Workloads with Mitigant

Attackers increasingly target AI workloads; however, as proven by NIST, there is no finite set of security controls that provides absolute security; the best guarantee is continuous AI red teaming. Determined adversaries continually look for means to break AI systems, and since these systems are constantly updated, coupled with evolving security issues and vulnerabilities, continuous testing with adversarial methods provides the best guarantee against attacks. 

‍

Across this two-part series, we discussed practical AI Red Teaming guidance for Amazon Bedrock core components. Knowing which attacks to run is only half the job; running them safely and repeatably in production workloads is ultimately critical. This is where Mitigant provides unbeatable value: by enabling security teams of any size to run AI Red Teaming for Amazon Bedrock continuously. Mitigant handles the entire attack lifecycle, including safe execution, attack analytics, cleanup, and comprehensive reporting. 

‍

Start your free trial at mitigant.io/sign-up or book a demo with our team.

‍

Übernehmen Sie die Kontrolle über Ihre Cloud-Sicherheitslage

Übernehmen Sie in wenigen Minuten die Kontrolle über Ihre Cloud-Sicherheit. Keine Kreditkarte erforderlich.