•
19 min read
On Secure Agent Architecture

An agent is more than an AI model responding to a prompt. It can maintain state, choose tools, execute actions, observe the results, and decide what to do next. Depending on its purpose, it may browse the web, write and run code, read email, query databases, modify files, operate cloud infrastructure, or communicate with other systems. This is what makes an agent useful, but also creates a different set of security challenges from a traditional application.

Traditional software executes control flow written by a developer. An agent generates part of its control flow at runtime. The model decides which tools to call, in what order, with what arguments, and when to stop. Those decisions are influenced by a context containing both trusted instructions and untrusted data.

A repository may contain a comment with hidden instructions to manipulate the agent. A web page may present data that looks like an instruction. An email, support ticket, document, or tool response may do the same. If the agent can take consequential actions, prompt injection is no longer only a problem with the quality of its answer. It becomes a path from untrusted content to real authority. A safer way to design around this is to arrange the system so that the model can never possess the authority it would need to cause unbounded harm. Therefore, a central security design principle I use throughout this post is that a model can propose actions, but its execution decision should be on deterministic infrastructure.

The central principle of this blog post is that a model can propose actions, but deterministic infrastructure decides whether they execute. The model is never responsible for its own security. The controls here run in code the model cannot edit, disable, or argue its way past. Prompt-level controls and model alignment are separate layers that I’ll cover in another post.

On Separate Reasoning, Authority, and Execution

An agentic system usually contains several things that may seem logical to collapse into one process:

  • The model and the loop that calls it
  • Session state and conversation history
  • Tools and generated code
  • Credentials
  • Network access
  • Policy and approval state
  • Logs and evidence

A lot of agent setups run the loop, shell/code execution, environment variables holding API keys, and outbound network in a single container. It means that anything escaping through one part of the system may reach everything else. Generated code can inspect the environment. A compromised tool can read session files. A prompt-injected agent can search for credentials. A malicious dependency can contact the network. A kill switch inside the same environment can be disabled by the thing it is supposed to stop.

Applying foundational systems-security principles to agents, we would want to separate the control plane from the data plane, and keep the component that enforces policy outside the component it’s enforcing against. In agent terms, it would separate what the system wants to do from what it’s permitted to do from what actually runs.

This means separation between three responsibilities.

The reasoning environment maintains the agent loop, context, plans, and durable session state. It decides what the agent would like to do.

The authority layer owns credentials, policy, external connections, approvals, budgets, and the ability to stop the system. It decides what the agent is permitted to do.

The execution environment runs generated code, shell commands, browsers, and other operations that the authority layer permits.

Reasoning, Authority, and Execution Layer

Deployments vary by context and organization. The reasoning environment may be a durable service while execution happens in short-lived microVMs. Policy, identity, and evidence may be distributed across several services.

The important properties that will always need to be true are:

  • Reasoning does not grant authority.
  • Authority is enforced outside the model.
  • Execution receives no ambient authority.
  • Evidence is recorded outside the environment performing the work.

On Reasoning

The reasoning environment gives the agent its adaptability. It maintains context, interprets observations, develops plans, and decides which action to propose next.

Treat Context as Mixed-Trust Data

An agent’s context can blend system instructions, operator policy, user requests, repository contents, web pages, emails, database records, tool output, messages from other agents, and the agent’s own earlier summaries.

Once inside the context window, these inputs all look the same. Telling the model which material is authoritative helps, but the security boundary should not rest on the model keeping that distinction. We need to be able to enforce it structurally. We should not have trusted policy live in writable workspace files that untrusted tools can modify. Durable session state shouldn’t live in disposable execution environments. And external content stays data, even when the model reads it as an instruction.

Preserve Provenance Through Memory and Summaries

Agent systems routinely summarize old context or store it in retrieval systems. As context gets transformed it becomes important to maintainance the provenance so that the trust level doesn’t change during later processing.

Memory records should retain their original source, ingestion time, producing agent or tool, transformation history, trust classification, related task, and supporting evidence.

Provenance gives later policy, agents, and reviewers a way to distinguish operator intent from conclusions derived from untrusted material.

Make Uncertainty Explicit

The reasoning environment should distinguish between different kinds of knowledge:

Status Meaning
Observed Captured directly from a tool or trusted system
Inferred Derived by the model from available evidence
Hypothesized A proposed explanation requiring testing
Verified Reproduced through an independent process

Model-generated output should not be stored as fact without verification in durable state. This is especially important when one agent’s output becomes another agent’s input. A model generated output should not acquire authority merely by crossing a message boundary or being stored in shared memory.

Keep Interpretation Separate from Evidence

An agent may misunderstand a tool response, describe a different command from the one that executed, or claim success after an operation failed. Its summaries, explanations, confidence assessments, and proposed next steps should remain separate from the actual action, policy decision, identity, command, result, and external side effects captured by trusted infrastructure.

This separation allows reasoning to remain flexible without making its account of events authoritative.

On Authority

The authority layer determines what the agent can cause to happen. It owns the identities, resources, policies, budgets, approvals, and external connections that give an action real effect. The model may request authority, but it should not create or widen it.

Make Actions Proposals

The agent should interact with the world by proposing structured actions to a trusted intermediary. A proposal describes the requested capability, target resource, intended operation, expected effects, reason for the request, and expected evidence or output.

The authority layer should then add everything the agent cannot assert for itself: authenticated identity, actual destination, resource ownership, data classification, budgets, approval state, rate limits, policy version, and session provenance.

It should then return one of three decisions:

Decision Meaning
Allow Perform the action with stated limits
Escalate Require human approval
Deny Do not perform the action

An action outside scope should not become permitted just because someone clicks an approval button. Human review is for exceptional actions within an authorized boundary.

Scope Capabilities

Policies expressed as ‘the agent may use Git, a browser, or a shell’ are too broad. A browser can read a page or submit a destructive form. Git can read a repository or rewrite history. A shell can run a parser or open a network tunnel.

The safer approach is to scope the capabilities:

repository.read
repository.write-branch
browser.navigate
browser.submit
database.read
database.mutate
artifact.publish
message.send
code.execute
identity.delegate

Capabilities should include resource and effect constraints. Permission to read one repository does not imply permission to read every repository. Permission to navigate a website does not imply permission to submit a form.

Every new tool creates a new action channel. It should not be enabled until its capabilities, policy path, resource limits, and audit behavior are defined.

Credentials Outside the Agent

An agent should never have access to a raw, long-lived secret, so a leaked context, transcript, or log can only disclose references. The agent should receive references to identities:

source-control-reader
test-environment-user
support-ticket-writer
production-observer

When an action is permitted, the authority layer should exchange the reference for a short-lived credential and use it for the specific operation. The secret never enters reasoning or execution. The permission granted for each action should be narrower than what the credential could do.

Credentials discovered during the agent’s work require the same treatment. Store the raw value in a protected system and return an opaque reference.

Bound Autonomy with Budgets

An action may be individually permitted and still become harmful through repetition. Agents can enter loops, retry failed operations, expand tasks recursively, create more agents, or continue gathering evidence after the answer is clear.

Useful budgets include tool calls, network requests, mutations, code executions, concurrent tasks, external messages, delegated agents, compute cost, wall-clock time, and repeated actions without new evidence.

Budgets belong in the authority layer. The model can know its remaining budget so that it can plan, but it should not be allowed to maintain or reset the counter. For multi-agent systems, budgets should apply to the delegation tree.

Human Review at Effect Boundaries

Human approval is most useful when an incorrect action would be difficult to reverse or would have a large blast radius.

Low-impact reads against disposable resources may run automatically. Sending an external message, modifying production state, publishing an artifact, accessing sensitive data, or performing an availability-risking operation may require approval or be prohibited.

Approval should bind to an exact action or tightly bounded plan in the following order: session, resource, identity, operation, parameters, expected effects, limits, and expiration.

A useful interface presents what will actually happen with the action beyond just the agent’s summary. Human gates should not compensate for weak deterministic policy; if reviewers must interpret every routine action, they will eventually approve by habit and the effectiveness of human review drops.

Kill Switch Outside the Agent

A kill switch belongs in the authority layer. It should stop forwarding actions, revoke delegated credentials, disable a capability, pause a task, stop an agent, stop a delegation subtree, or stop the entire session.

Stopping new actions at the broker should happen before waiting for credential revocation or sandbox termination. Preserve evidence before destroying the execution environment.

On Execution

The execution environment is where permitted actions are executed. It runs generated code, shell commands, browsers, test runners, and other tools.

Treat the Execution Environment as Untrusted

Treat the execution environment as compromised from the moment it starts.

The boundary should sit below the language runtime. Import restrictions, object filters, and model-based code review may reduce accidental misuse, but they are not dependable isolation for hostile code.

Plain containers share the host kernel. User-space kernels such as gVisor place another boundary between the workload and host. MicroVMs such as Firecracker or Kata give the workload its own kernel while remaining fast enough to create per task or session. Full VMs or dedicated hosts give the strongest separation at higher cost and startup time; use them where the risk justifies it.

A secure execution environment should normally have no credentials or cloud identity, no writable host mounts, a read-only base image, a disposable workspace, no unrestricted network route, strict resource limits, external monitoring, and a termination mechanism controlled from outside.

The base image itself is part of the containment boundary. Use a pre-verified base with no package manager or registry access, and never bake secrets into a shared image layer.

Assume the sandbox will be escaped. Design so that a compromised sandbox holds as little enduring authority as possible.

Default-Deny the Network

Unrestricted outbound access turns local compromises into lateral movement. Malicious code can download tools, upload data, contact command-and-control infrastructure, probe internal services, reach metadata endpoints, or publish modified artifacts.

The execution environment should have no default route to the internet or internal infrastructure. Required traffic should pass through a broker or proxy governed by the authority layer.

Network permission should be specific to destination, protocol, method, resource, identity, data class, request size, rate, and time window.

DNS requires the same treatment. A trusted resolver can resolve approved resources, validate the resulting address, and bind resolution to the permitted connection. Redirects should be treated as new destinations.

Browser-based agents need controls over scripts, images, WebSockets, prefetching, service workers, redirects, and downloads

The model API endpoints may also be an egress path. Provider-side web search, connectors, file retrieval, and URL fetching should be disabled or governed by the same destination policy.

Ephemeral Execution

Long-lived execution environments accumulate generated code, downloaded files, modified tools, caches, untrusted content, and state whose provenance becomes difficult to reconstruct.

Create execution environments for bounded tasks or sessions. Impose hard limits, preserve required evidence, and then destroy them.

The logical agent does not have to be ephemeral. Its durable state can survive outside the worker.

Treat the Lifecycle as Part of the Boundary

Agents wait for input, suspend while idle, resume on demand, move between workers, and divide work into child tasks. These lifecycle transitions can change the relationship between state, identity, policy, and execution.

logical agent
     ├── durable state
     ├── task and delegation history
     └── current policy context
              │ activate
              ▼
       disposable worker
       ├── process tree
       ├── temporary filesystem
       └── short-lived identity

Trusted infrastructure should verify that state belongs to the expected agent, the runtime has passed integrity checks, the worker is currently assigned, old state is gone, current policy is installed, credentials are bound to this activation, budgets and revocations are current, and external logging is available.

Treat Snapshots as Executable State

A snapshot may contain process memory, filesystem changes, model context, untrusted content, generated code, stale credentials, or persistence created during a compromise.

Snapshots should be encrypted, integrity-protected, bound to one logical agent and runtime version, inaccessible to the sandbox, checked against current revocation state, and verified before restoration.

A snapshot captured after suspicious behavior may remain useful for forensics, but it should be marked as tainted and prevented from automatically returning to service.

Treat Worker Reuse as a Tenant Transition

Assigning a reused worker to another agent is not merely starting a new process. It is changing tenants.

The platform must reset process state, memory, filesystem overlays, environment variables, mounts, network policy, local sockets, credentials, caches, and tracing context.

This transition should be tested directly with canary state across suspend, resume, migration, and reassignment.

Cross-Cutting Controls

Some architectural elements span reasoning, authority, and execution.

Record Actions Outside the Agent

Capture proposed actions, policy decisions, resolved identities and resources, execution starts, actual requests or commands, results, side effects, approvals, and tool and policy versions at the enforcement point.

Store those records somewhere reasoning and execution cannot modify. Separate each operation into at least:

proposed
started
completed

Evidence should follow the logical agent rather than the physical worker. Logs and traces should carry agent, task, activation, parent task, snapshot, policy, and worker-assignment identifiers.

Detect and Respond

Detection should be designed to consume the enforcement-point record, and produce decisions for the authority layer to execute.

The high-value signals are generally behavioral such as, a spike in denials (an agent probing its boundary), a move from read to mutation without approval, a discovered credential followed by an outbound request, loop depth, budget burn rate, and in multi-agent systems an edge the topology did not declare or a message whose provenance is too low for what it would trigger. Single-event rules catch some. Chain rules, an ordered sequence within a bounded window, catch what looks normal one step at a time.

Response has two tiers, a circuit breaker that throttles or suspends on a threshold, and a kill switch that revokes and destroys while preserving evidence first. Most anomalies require the circuit breaker. Live alerting and retrospective review should run the same rule definitions, and should not disagree about what is anomalous.

Detection is probabilistic and acts after the fact, so it should not be the boundary for an irreversible action. Detection exists to catch what policy did not foresee. And under heavy class imbalance it runs precision-first, which means accepting misses to avoid drowning reviewers. A human review still matters the most when intent is the only thing that differentiates a malicious actor from a legitimate work.

Treat Agent Relationships as Boundaries

Messages, task assignments, shared memory, artifact stores, queues, and delegation allow one agent to influence another. Our system design should treat every transfer as a mediated edge.

The sending agent provides a payload, while trusted infrastructure should attach verified sender and recipient, capabilities, delegation chain, resource scope, provenance, and correlation identifiers.

In multi-agent systems, the following three invariants act as important controls to protect the boundary:

Delegation attenuates. A hop in the chain can only narrow authority. If an orchestrator agent delegates to a worker agent, the worker’s capabilities are a subset of the orchestrator’s. A compromised worker cannot unilaterally escalate itself; it can only act within what was explicitly delegated down.

Provenance floors persist across hops. Untrusted data should stay untrusted always. If a message originated from low-trust input, that floor travels with it through every agent that touches it. One clean relay does not erase the source.

Budgets apply to the whole team. A single budget ceiling applies across all agents in a delegation hierarchy. Spreading work across many child agents should not be able to beat the ceiling by splitting charges across them. This needs to be enforced by the authority layer.

Architecture Checklist

  1. Are reasoning, authority, and execution separated by explicit trust boundaries?
  2. Is model output treated as a proposal rather than a source of authority?
  3. Does durable memory retain source provenance and uncertainty?
  4. Does every consequential action cross an enforcement point?
  5. Does policy depend on authority-layer facts rather than agent assertions?
  6. Are capabilities scoped by resource and effect rather than only by tool?
  7. Are real credentials kept outside reasoning and execution environments?
  8. Are budgets enforced outside the model?
  9. Are human approvals bound to exact actions or constrained plans?
  10. Can the authority layer stop actions without agent cooperation?
  11. Does untrusted code execute below a hardened isolation boundary?
  12. Does execution begin without credentials or ambient cloud identity?
  13. Is outbound network access denied by default?
  14. Are DNS, redirects, browser subresources, and provider-side retrieval governed?
  15. Is execution disposable even when the logical agent is durable?
  16. Do lifecycle transitions re-establish identity, policy, budgets, and logging?
  17. Are snapshots protected and verified before restoration?
  18. Is worker state reset between agent assignments?
  19. Is evidence captured outside reasoning and execution?
  20. Can delegation only reduce authority?
  21. Are shared stores treated as agent-to-agent communication edges?
  22. Do logs follow the logical agent across workers and activations?
  23. Can the authorization path be explained without trusting the model’s judgment?
  24. Does an alert in the evidence stream trigger response without waiting for a human?

What This Architecture Does Not Solve

The design patterns proposed in this blog removes ambient authority and enforces boundaries deterministically. It does not solve whether an agent, operating within its constraints, is pursuing the right objective.

  • An agent may still misunderstand its task, miss important information, produce poor plans, or remain inside its formal permissions while pursuing the wrong objective.
  • Deterministic policy can enforce resources, identities, operations, rates, and budgets. It cannot fully determine whether an action is semantically relevant to the user’s intent.
  • Provenance remains imperfect. Infrastructure can identify where information entered the system, but cannot reliably determine how strongly a later conclusion depends on a poisoned input.
  • The agent’s output may itself become an exfiltration channel. Restricting what it can read is generally stronger than trying to recognize every possible encoding in its output.
  • Human reviewers may still approve bad actions. Multiple agents may coordinate through permitted messages, shared timing, target state, or other channels the architecture does not completely model.
  • The authority layer also becomes critical infrastructure. It should remain smaller, simpler, and more heavily reviewed than the system it controls. If every tool implements its own exceptions, the enforcement point becomes a collection of inconsistent bypasses.

Conclusion

One of the most important security decision in an agent architecture is where authority lives. If reasoning, execution, credentials, policy, and network access all live inside one environment, a failure in any one of them can become a failure of the whole system.

Separating them can provide structural protections:

  • The model can propose without authorizing.
  • Generated code can execute without holding credentials.
  • Credentials can be used without entering the sandbox.
  • External actions can be mediated without trusting model judgment.
  • Evidence can survive a compromised workspace.
  • Durable state can survive without making workers durable.
  • The system can stop without asking the agent to cooperate.

Strong isolation, restricted egress, scoped capabilities, budgets, external monitoring, and a kill switch still matter. They become easier to reason about once the architecture is designed around keeping deterministic authority on top of probabilistic model reasoning.