AI is a general-purpose technology, like electricity, internet and transport. We don’t slow down their development or try to avoid them; we just regulate, ensure their safety, and prosecute misuse.
Recent stories about AI agents “hacking public sites” or “escaping sandboxes” are easy to frame as an AI-control problem. Usually, the more useful technical diagnosis is simpler: a system gave untrusted automation a path to influence something more trusted.
That is serious. It deserves patches, careful deployment, and incident response. But it is not evidence that a model has acquired independent intentions. It is evidence that agentic systems must be designed like any other security-sensitive automation: with small privileges, strong boundaries, explicit approvals, and useful audit trails.
Thesis: Treat an AI agent as capable, untrusted automation—not as a harmless chatbot, and not as an autonomous adversary with mysterious powers.
What an “escape” often means
A sandbox can be intact while the whole system remains unsafe. A common failure chain looks like this:
Agent in sandbox
-> writes repository configuration / generated artifact / task output
-> trusted host tool reads that output
-> host tool executes it with broader permissions
-> agent-controlled data crosses a trust boundary
The defect is not necessarily a broken container boundary. It may be a trust handoff: a host-side component treats agent-controlled files or output as trusted configuration.
Untrusted agent output
|
v
[validator / policy gateway] ---- reject or require review
|
v
[approved, narrow tool call] ---- scoped identity + audit log
|
v
[isolated target environment]
This is familiar application security. It resembles unsafe deserialization, CI/CD supply-chain compromise, malicious Git hooks, or an unvalidated upload later processed by a privileged service.
A better framing
When somebody says “the agent went out of control,” we must ask four engineering questions:
- What capability did it actually have? File write, shell execution, network access, credentials, deployment rights, or all of them?
- Which trust boundary failed? The container, a host integration, an API authorization check, a configuration loader, or an overly broad token?
- Which component interpreted agent-controlled data? A shell, IDE extension, CI runner, package manager, Git hook, or deployment tool?
- What is the blast radius? One temporary worktree, one developer machine, one staging project, or production?
Those questions move the discussion from vague fears about intelligence to testable properties of a deployed system.
Important distinction: A model can produce an effective exploit, or optimize toward task completion, without possessing consciousness, personal motives, or independent political goals. Security controls should rely on observed capabilities and permissions, not assumptions about a model’s inner state.
Threat model: the agent is a semi-trusted principal
A practical default is to model the agent as an internal user whose output may be adversarial, whether due to prompt injection, compromised dependencies, mistaken instructions, or an operator error.
| Asset | Unsafe default | Safer default |
|---|---|---|
| Filesystem | Agent sees the whole repository and home directory | Dedicated worktree with explicit writable paths |
| Network | Unrestricted outbound internet access | Deny by default; allow-list required APIs |
| Credentials | Shared developer or CI token | Per-session, short-lived, scoped token |
| Shell | Arbitrary command execution | Allow-listed commands or approval-gated execution |
| Deployment | Agent can push or deploy directly | Agent opens a PR; CI and human approval deploy |
| Configuration | Agent can edit hooks and tool settings | Treat configuration changes as privileged code review |
Secure-by-default patterns
Use separate identities
Do not run an agent as your normal developer account or as a broad CI service account. Give each run a distinct identity whose permissions expire with the task.
Bad:
Agent -> developer's SSH key, cloud credentials, package tokens
Better:
Agent -> temporary task identity
-> read/write: ./workspace only
-> network: api.example.internal only
-> expires: 30 minutes
-> cannot deploy, rotate secrets, or alter IAM
The goal is not to make an agent “trustworthy.” The goal is to make mistakes and compromised prompts cheap to contain.
Separate planning from execution
An agent can propose commands and patches without automatically receiving the authority to execute them.
# conceptual agent policy
permissions:
read:
- workspace/**
write:
- workspace/**
execute:
allowed_commands:
- pytest
- ruff
- npm test
- git diff
network:
allow:
- registry.npmjs.org
privileged_actions:
require_human_approval:
- git push
- docker build
- kubectl
- terraform
- any_secret_access
The exact syntax varies by product. The design principle does not: tool access is an authorization problem, not a prompt-writing problem.
Keep outputs as data
Never silently treat agent-generated content as executable configuration. Review or validate any generated:
- Shell scripts and command strings
- CI/CD workflow files
- Git hooks and repository metadata
- Dockerfiles and container entrypoints
- Package-manager configuration
- IDE task definitions and extensions
- Infrastructure-as-code files
# Bad: generated text becomes a shell program
subprocess.run(agent_output, shell=True, check=True)
# Better: select only approved operations and pass arguments as data
ALLOWED = {
"test": ["pytest", "-q"],
"lint": ["ruff", "check", "."],
}
subprocess.run(ALLOWED[requested_action], check=True, cwd="workspace")
The safer version still needs authorization around requested_action, but it removes a large class of command-injection risk.
Use a real egress policy
A container without network policy is not a complete sandbox. Restrict both where an agent can connect and which credentials it can reach.
Default network policy:
ingress: deny
egress: deny
Task exception:
allow HTTPS only to:
- package registry mirror
- internal documentation proxy
- explicitly selected test service
Never expose:
- cloud instance metadata services
- local Docker socket
- host loopback administrative APIs
- unrestricted DNS resolvers
A strong boundary should survive hostile text in an issue, repository, web page, or document. Prompt instructions are not a security boundary.
Make deployment deliberately boring
For most teams, an agent should produce a branch or patch—not alter production.
Agent -> isolated worktree -> commits patch -> opens pull request
-> CI runs tests and security checks
-> human review and protected-branch rules
-> deployment pipeline uses a separate identity
This preserves the useful part of coding agents while keeping high-impact actions behind ordinary software-delivery controls.
Minimum checklist
Use this before enabling an agent in a repository, CI job, or internal application.
Identity and permissions
- Each task runs with a separate, short-lived identity
- The agent has only the filesystem paths it needs
- Credentials are scoped, short-lived, and never placed in prompts or persistent memory
- The agent cannot access the developer’s SSH keys, browser profile, cloud CLI credentials, or Docker socket
- Production, IAM, billing, secret rotation, and deletion permissions are unavailable by default
Isolation and network
- The agent runs in a disposable container, microVM, or equivalent isolated environment
- Host mounts are read-only or absent unless explicitly needed
- Outbound network access is denied by default and allow-listed per task
- Access to metadata endpoints, host loopback services, and local administrative APIs is blocked
- The environment is discarded after the task rather than reused indefinitely
Tools and approvals
- Tools are individually authorized; “shell access” is not the default permission
- Destructive or external actions require a human confirmation step
- An agent cannot directly deploy, publish packages, merge protected branches, or modify access policies
- Generated configuration, hooks, workflows, and infrastructure changes receive code review
- Tool arguments are structured and validated rather than built from free-form model text
Observability and response
- Prompts, tool calls, outputs, changed files, and external requests are logged with task identity
- Logs do not store plaintext secrets or private user data unnecessarily
- Alerts exist for permission failures, unusual egress, new destinations, and privilege-escalation attempts
- The team can revoke the agent identity and terminate its runtime quickly
- Prompt injection and malicious-repository scenarios are included in regular security tests
What not to conclude
We should not dismiss the risks. Agents can lower the cost of phishing, reconnaissance, insecure code changes, and abuse of a badly designed integration. They can also make ordinary security mistakes faster and more scalable.
But we also should not treat every sandbox failure as proof that artificial systems have become independent actors. A coding agent with unbounded shell access, credentials, writable configuration, and unrestricted network egress is not a reliable safety experiment; it is an over-privileged automation service.
The practical response is familiar:
- Reduce privilege
- Validate boundary crossings
- Separate duties
- Require approval for irreversible actions
- Log behavior
- Limit blast radius
- Test adversarially
That approach is less dramatic than an “AI rebellion” story. It is also more actionable—and it makes systems safer whether the next harmful action comes from an AI agent, a compromised dependency, a malicious insider, or a human mistake.
References:
- OWASP, AI Agent Security Cheat Sheet: per-tool authorization, trust separation, and agent-specific security controls. OWASP
- NIST, Agentic AI: Emerging Threats, Mitigations and Challenges: scoped tools, sandboxing, signed goals, and approval mechanisms. NIST
- Docker, How to Secure AI Agents: isolation, tool access, identity, and runtime monitoring. Docker