Locking Down AI Agent Tool Permissions in the Cloud

Securing cloud AI agents is not a prompt-engineering problem, and AI agent tool permissions are exactly where the risk concentrates. The controls that survive production are a deny-first tool permission layer, an execution identity scoped to a single agent or workload, and an approval gate on every third-party tool the agent can reach. The model’s own judgement is not the boundary: in Claude Code and the Claude Agent SDK, permission rules are evaluated in a fixed order — deny, then ask, then allow — and a matching deny blocks the call even when a narrower allow rule also matches. That precedence, combined with least-privilege IAM on the cloud side and deliberate vetting of MCP tool servers, is what actually constrains what an agent can touch.

The Threat Model for Agent Tools

Most successful attacks against agentic systems do not break cryptography; they abuse the metadata that travels with tools. Tool poisoning hides instructions inside a tool description or code comment, invisible to the user but fully visible to the model, steering the agent toward reading secrets or calling endpoints it should never touch. Rug pulls are worse: a server behaves correctly at review time, then changes its tool descriptions after approval, and because the protocol does not prohibit later modification, clients keep invoking a tool whose behaviour has silently changed. Installations that pull the server fresh on every run, as npx and uvx do, amplify this because there is no pinned artifact to audit.

The measured impact is not theoretical. In an end-to-end empirical evaluation of the MCP ecosystem, exploitation via malicious external resources reached a 93.33% attack success rate, with tool poisoning attacks at 59.33%. The same study uploaded malicious servers to every aggregation platform it tested without rejection, and concluded that current mainstream models cannot reliably defend against these vectors on their own. For engineering teams the lesson is blunt: the model is the last line of defense, never the first.

Deny-First Rules Win

The permission layer documented for Claude Code and the Agent SDK is the pattern most platforms should copy. Rules are declarative and evaluated with strict precedence: a bare-name deny such as Bash removes the tool from the model’s context entirely, so the model cannot even attempt the call; a scoped deny such as Bash(curl *) keeps the tool available but blocks matching invocations in every permission mode, including bypass modes. Deny rules set in managed settings cannot be overridden by developers, command-line flags or project files — if a tool is denied at any level, no other level can allow it. Hooks run before the approval prompt and can deny or force a prompt, but a hook can never loosen a deny rule.

RuleEffectExample
Deny, bare tool nameRemoves the tool from the model’s context entirelydeny: Bash
Deny, scoped patternBlocks matching calls in every mode, including bypassdeny: Bash(curl *)
AskPauses for human confirmation before the call runsask: WebFetch(domain:*)
AllowAuto-approves the matched call; unlisted tools fall through to the active modeallow: Read(docs/**)

For an unattended cloud job, pair an explicit allow list with a non-interactive deny-by-default mode: listed tools run, everything else is refused without prompting. That fail-closed combination is the right default when no human is watching the loop.

Scoping IAM Roles on AWS

On AWS the agent’s permissions are IAM, and the pattern AWS itself documents is narrow enough to reuse. A Bedrock agent runs under a service role scoped to the specific foundation models it may invoke, the S3 objects holding its action group schemas, and the knowledge bases it queries — nothing broader. To force inference through the agent’s enforced configuration, the documented pattern allows a role to invoke one specific agent alias while direct bedrock:InvokeModel calls are denied outright, so guardrails, approved prompts and KMS settings applied to the agent cannot be sidestepped by calling the raw model API.

Every action group Lambda function needs a resource-based policy that lets the agent’s service role invoke it — the permission lives on the function, not only in the role, and both must exist for the agent to act. Keep the caller identity, the assumed role and the execution identity separate, and grant each only what its own task requires. Deployment credentials should never be reachable from the environment where commands and MCP tools execute.

Vetting MCP Tools Before Deployment

Third-party tool servers deserve their own gate. On Anthropic’s managed agents API, MCP toolsets default to always_ask so that tools added to a server later cannot run without approval — a direct defense against rug pulls, and a default worth keeping before you override it with blanket auto-approval. Pin server versions to reviewed artifacts instead of resolving the latest tag at every start, diff tool descriptions whenever a server updates, and treat every tool output and fetched page as untrusted input rather than instruction. Layer an egress allowlist underneath: even a fully poisoned tool cannot exfiltrate what it cannot send. A fail-closed approval gate beats a clever prompt every time, because the gate does not depend on the model noticing it is under attack.

A Hardening Checklist

Run this before an agent touches anything with a credential attached:

  1. Inventory every tool, MCP server and API the agent can reach, including tools injected transitively by dependencies.
  2. Write the deny layer first: bare-name denies for anything interactive, scoped denies for destructive commands, then ask rules for the grey zone, then a minimal allow list.
  3. Pair the allow list with a non-interactive deny-by-default mode for unattended jobs.
  4. Scope the execution role to one agent and its named resources, and block direct model invocation outside the agent path.
  5. Attach resource-based policies to every action group function and verify the trust path end to end in a staging account.
  6. Pin and diff MCP tool definitions, and re-review whenever a description or version changes.
  7. Enforce egress allowlists so both approved tools and compromised tools can only reach known endpoints.

Permission hygiene also pays off on cost: the same review that decides which tools an agent may call should revisit how prompt caching changes request economics and which model routing choices are worth automating, since both determine how many tool-calling loops a workload can afford.

Sources