Skip to content
Radoslav
Go back

We should worry about how humans deploy AI, not about AI spontaneously deciding to harm us.

Updated:
Edit page

AI is a general-purpose technology, like electricity, internet and transport. We don’t slow down their development or try to avoid them; we just regulate, ensure their safety, and prosecute misuse.

Recent stories about AI agents “hacking public sites” or “escaping sandboxes” are easy to frame as an AI-control problem. Usually, the more useful technical diagnosis is simpler: a system gave untrusted automation a path to influence something more trusted.

That is serious. It deserves patches, careful deployment, and incident response. But it is not evidence that a model has acquired independent intentions. It is evidence that agentic systems must be designed like any other security-sensitive automation: with small privileges, strong boundaries, explicit approvals, and useful audit trails.

Thesis: Treat an AI agent as capable, untrusted automation—not as a harmless chatbot, and not as an autonomous adversary with mysterious powers.

What an “escape” often means

A sandbox can be intact while the whole system remains unsafe. A common failure chain looks like this:

Agent in sandbox
  -> writes repository configuration / generated artifact / task output
  -> trusted host tool reads that output
  -> host tool executes it with broader permissions
  -> agent-controlled data crosses a trust boundary

The defect is not necessarily a broken container boundary. It may be a trust handoff: a host-side component treats agent-controlled files or output as trusted configuration.

Untrusted agent output
        |
        v
[validator / policy gateway] ---- reject or require review
        |
        v
[approved, narrow tool call] ---- scoped identity + audit log
        |
        v
[isolated target environment]

This is familiar application security. It resembles unsafe deserialization, CI/CD supply-chain compromise, malicious Git hooks, or an unvalidated upload later processed by a privileged service.

A better framing

When somebody says “the agent went out of control,” we must ask four engineering questions:

  1. What capability did it actually have? File write, shell execution, network access, credentials, deployment rights, or all of them?
  2. Which trust boundary failed? The container, a host integration, an API authorization check, a configuration loader, or an overly broad token?
  3. Which component interpreted agent-controlled data? A shell, IDE extension, CI runner, package manager, Git hook, or deployment tool?
  4. What is the blast radius? One temporary worktree, one developer machine, one staging project, or production?

Those questions move the discussion from vague fears about intelligence to testable properties of a deployed system.

Important distinction: A model can produce an effective exploit, or optimize toward task completion, without possessing consciousness, personal motives, or independent political goals. Security controls should rely on observed capabilities and permissions, not assumptions about a model’s inner state.

Threat model: the agent is a semi-trusted principal

A practical default is to model the agent as an internal user whose output may be adversarial, whether due to prompt injection, compromised dependencies, mistaken instructions, or an operator error.

AssetUnsafe defaultSafer default
FilesystemAgent sees the whole repository and home directoryDedicated worktree with explicit writable paths
NetworkUnrestricted outbound internet accessDeny by default; allow-list required APIs
CredentialsShared developer or CI tokenPer-session, short-lived, scoped token
ShellArbitrary command executionAllow-listed commands or approval-gated execution
DeploymentAgent can push or deploy directlyAgent opens a PR; CI and human approval deploy
ConfigurationAgent can edit hooks and tool settingsTreat configuration changes as privileged code review

Secure-by-default patterns

Use separate identities

Do not run an agent as your normal developer account or as a broad CI service account. Give each run a distinct identity whose permissions expire with the task.

Bad:
  Agent -> developer's SSH key, cloud credentials, package tokens

Better:
  Agent -> temporary task identity
        -> read/write: ./workspace only
        -> network: api.example.internal only
        -> expires: 30 minutes
        -> cannot deploy, rotate secrets, or alter IAM

The goal is not to make an agent “trustworthy.” The goal is to make mistakes and compromised prompts cheap to contain.

Separate planning from execution

An agent can propose commands and patches without automatically receiving the authority to execute them.

# conceptual agent policy
permissions:
  read:
    - workspace/**
  write:
    - workspace/**
  execute:
    allowed_commands:
      - pytest
      - ruff
      - npm test
      - git diff
  network:
    allow:
      - registry.npmjs.org
  privileged_actions:
    require_human_approval:
      - git push
      - docker build
      - kubectl
      - terraform
      - any_secret_access

The exact syntax varies by product. The design principle does not: tool access is an authorization problem, not a prompt-writing problem.

Keep outputs as data

Never silently treat agent-generated content as executable configuration. Review or validate any generated:

# Bad: generated text becomes a shell program
subprocess.run(agent_output, shell=True, check=True)

# Better: select only approved operations and pass arguments as data
ALLOWED = {
    "test": ["pytest", "-q"],
    "lint": ["ruff", "check", "."],
}

subprocess.run(ALLOWED[requested_action], check=True, cwd="workspace")

The safer version still needs authorization around requested_action, but it removes a large class of command-injection risk.

Use a real egress policy

A container without network policy is not a complete sandbox. Restrict both where an agent can connect and which credentials it can reach.

Default network policy:
  ingress: deny
  egress: deny

Task exception:
  allow HTTPS only to:
    - package registry mirror
    - internal documentation proxy
    - explicitly selected test service

Never expose:
  - cloud instance metadata services
  - local Docker socket
  - host loopback administrative APIs
  - unrestricted DNS resolvers

A strong boundary should survive hostile text in an issue, repository, web page, or document. Prompt instructions are not a security boundary.

Make deployment deliberately boring

For most teams, an agent should produce a branch or patch—not alter production.

Agent -> isolated worktree -> commits patch -> opens pull request
      -> CI runs tests and security checks
      -> human review and protected-branch rules
      -> deployment pipeline uses a separate identity

This preserves the useful part of coding agents while keeping high-impact actions behind ordinary software-delivery controls.

Minimum checklist

Use this before enabling an agent in a repository, CI job, or internal application.

Identity and permissions

Isolation and network

Tools and approvals

Observability and response

What not to conclude

We should not dismiss the risks. Agents can lower the cost of phishing, reconnaissance, insecure code changes, and abuse of a badly designed integration. They can also make ordinary security mistakes faster and more scalable.

But we also should not treat every sandbox failure as proof that artificial systems have become independent actors. A coding agent with unbounded shell access, credentials, writable configuration, and unrestricted network egress is not a reliable safety experiment; it is an over-privileged automation service.

The practical response is familiar:

That approach is less dramatic than an “AI rebellion” story. It is also more actionable—and it makes systems safer whether the next harmful action comes from an AI agent, a compromised dependency, a malicious insider, or a human mistake.

References:


Edit page
Share this post:

Next Post
Scalar, Matrix and Tensor