If your agent needs to run Python over a customer’s CSV, the code has to execute somewhere. Either you operate a sandbox — an E2B, Modal, or a Docker container you keep patched, isolated, and time-capped — or you add one entry to the tools array and let the provider run commands on its own infrastructure. OpenRouter’s comparison of server-side code execution tools, published October 5, 2026, lays out what each option actually involves, and the tradeoffs are worth internalizing before you pick.
How server-side execution actually works
A model never runs anything itself. It emits a tool call; something has to carry it out. With a client-side tool, your application receives the call, executes it, and sends results back in a follow-up request. With a server-side tool, the provider runs the command in its own sandbox and returns stdout, stderr, and an exit code (or timeout) to the model within the same request.
The value is in the loop: the model runs a command, reads real output, decides whether to run another, and repeats until it answers. OpenRouter caps that loop with a max_tool_calls parameter, defaulting to 30 steps per request. Your app never writes a handler for any of it.
Four providers, one key difference
OpenAI, Anthropic, and Google each run code only for their own models. OpenAI’s shell tool works on the Responses API, runs Debian 12 without sudo, and supports network access only after an admin configures an allowlist. Anthropic’s code execution tool gives Claude Python and Bash on the Messages API — 5 GiB RAM, one CPU, no outbound connections, so no installing packages mid-run. Google’s tool for Gemini is Python-only with a 30-second runtime cap and a fixed library list.
OpenRouter’s tools — openrouter:shell (Responses and Messages APIs) and openrouter:bash (Messages only, both in beta) — run at the routing layer, so they work with any model on those APIs. The practical consequence: if you switch models, the tool call in your request doesn’t change. The engine parameter set to auto keeps a provider’s native shell where one exists and falls back to OpenRouter’s sandbox otherwise. For openrouter:bash, auto returns the call to your app to run client-side instead.
Billing is time-based: $0.0001 per second with a 30-second minimum for a new or sleeping container. Sandboxes are scoped to your account and workspace, with outbound network off by default.
We covered the shape of this earlier in provider-run sandboxes becoming part of the API; what this comparison adds is the concrete terms of the trade.
What you hand over, and what you keep
The comparison’s most useful section is what it says about the self-managed column. A hosted tool trades control for a JSON entry: you don’t size the container, patch the image, or own the security boundary. The costs show up elsewhere.
- Runtimes shift. OpenRouter’s sandbox reported Ubuntu 22.04.5 with Python 3.11.14 when they tested on September 22, 2026, and they explicitly warn the image can change — read versions from the container rather than hard-coding them.
- Latency and cost are per second. Short commands are cheap; a chatty loop that burns through 30 tool-call steps adds up.
- Environments are fixed. No custom base image, no GPU, no long-running sessions. All four providers’ hosted tools fall short on those, and that’s the honest line: if your workload needs any of them, run your own sandbox via a platform like E2B or Modal.
Hosted execution fits bounded work — running a script, transforming a file, checking a result against real output. If your agent loops, also think about evaluation: when the sandbox image or model changes underneath you, your agent’s behavior changes too, so treating tool-loop regressions seriously matters.
The practical next step: pick one bounded task in your agent, wire in a hosted shell tool, and measure seconds billed and steps used per request. If the numbers work, you’ve deleted a container you never have to patch.
