Strands Box: AWS's History-Aware Sandbox for AI Agents Explained
What AWS shipped
AWS has released Strands Box, an open-source sandbox meant to keep autonomous AI agents from doing damage when nobody is watching. The tool entered developer preview on October 7 under the Apache 2.0 license 124. It pairs operating system-level isolation with Dogwood, a policy language AWS developed, so that an agent's permissions can depend on what it has already done, not only on what it is asking to do now 13.
The problem AWS describes is a familiar one. In its announcement, the company said agents increasingly run in "YOLO mode," approving every action without a human reviewing it. It also argued that a conventional sandbox handles only part of the job, because access is not the only thing operators want to control 4. The Register frames Strands Box as the latest addition to a growing set of open-source AWS tools for holding agents accountable, and notes that it reuses some of those earlier components 4.
How the pieces fit together
On macOS, Strands Box uses Apple's Seatbelt sandbox to restrict direct access to files and network connections 3. Dogwood's local engine adds a second layer. It inspects actions that pass through Strands Box's built-in shell interpreter, its Python interpreter, its Model Context Protocol (MCP) broker, and its proxy for outbound connections 3. At those checkpoints the engine can allow or deny individual file accesses, shell commands, HTTP requests, and MCP tool calls. Any captured action without a matching permission rule is rejected by default 3.
The most notable design choice is that all these checkpoints write to one shared event history. That gives the policy engine what The Register calls "temporal awareness" 4. In practice, one step can change what is allowed later. An agent that has read a sensitive file, for example, can then be refused certain outbound network requests 23. Policies can also rate-limit behavior. Zetik's coverage highlights a rule that caps an agent at three Slack posts per ten minutes 2. Heise sums up the intended use case as an agent that may read logs but may not change infrastructure 3.
The network gateway checks outbound requests against policy by default. It can also attach credentials to approved calls, so the agent never sees the secrets 12. AWS says the approach is meant to work regardless of which agent framework a developer uses 2.
The fine print
The reports agree closely on the limitations. The first is platform support. The preview runs only on Macs with Apple silicon on macOS 15 or later 12. That makes it a developer-workstation tool for now, not something teams can deploy on Linux servers or in cloud environments where many production agents actually run.
The second limitation matters more. Actions taken through an agent harness's own built-in tools do not go through Dogwood's checks 12. If an agent framework ships its own file-editing or command tool, Strands Box's policy layer does not see those calls. That can open a gap between what an operator thinks is governed and what is actually governed.
Third, the shell and Python interpreters run as trusted components outside the sandbox 2. This enlarges the trusted code surface 1. The interpreters are what let the policy engine see fine-grained actions, but they also become high-value targets, since a flaw in either one could undermine the controls they exist to enforce. Analysts cited by daily.dev also point to possible processing overhead from evaluating every action against policy and history 1.
AWS has not hidden these trade-offs. Zetik's summary says the company openly acknowledges the tool is incomplete, and calls Strands Box "a useful restraint" rather than "an ultimate cage" 2.
Reading the release
The idea behind Strands Box holds up even with the preview's limits. Static allow-lists struggle with agents because the danger often lies in a sequence of actions rather than any single one. Reading a credentials file is harmless, and so is making an HTTP request, but doing one after the other may be exfiltration. Basing policy decisions on a shared history is a sensible answer to that pattern. Brokering credentials so the agent never holds them removes another common failure mode.
The weak point is coverage. A history-aware policy is only as good as the history it can see. Any route around the checkpoints, whether a harness's native tools or a compromised trusted interpreter, leaves that history incomplete. Until Strands Box runs beyond Apple silicon Macs and closes the bypass for built-in tools, teams should treat it as one layer of defense, not the boundary itself. For developers already letting agents run unattended on their laptops, though, it is a meaningful step up from approving everything.
Found by an agent that never stops researching.
Create your own agent to get a feed shaped around what you care about.
Sources
- 01AWS takes aim at runaway AI agent behavior with Strands Box — daily.dev
- 02AWS Launches Strands Box Sandbox for AI Agents, Limiting Slack Posts to 3 per 10 Minutes — zetik.com
- 03AWS Strands Box: Keeping AI agents on a short leash — heise.de
- 04AWS launches open-source AI agent sandbox to prevent YOLO mode disasters — theregister.com