If your agent needs to run Python over an uploaded CSV, the first real question isn’t which model to use — it’s where the code runs. Run it yourself and you’re picking a sandbox platform like E2B or Modal, or standing up your own container, then handling isolation from your database, runtime caps, and image patching. Run it through the provider and it’s one entry in the request’s tools array. OpenRouter published a comparison of four hosted code execution tools on October 5, 2026, and it’s a useful map of how this decision is settling.
The pattern, in one request
Server-side code execution is the same trick as hosted web search: the model emits a tool call, the provider carries it out on its own infrastructure, and the result goes back to the model inside the same request. Your application has no handler for it. The model runs a command, reads stdout and stderr, and decides whether to run another — and that loop is the whole value, because the model checks its work against real output instead of guessing.
The loop isn’t unlimited. OpenRouter’s max_tool_calls field caps server-tool steps per request, with a default and maximum of 30.
Four providers, four scopes
The dividing line between the hosted tools is which models they serve:
- OpenAI runs a shell on the Responses API for OpenAI models — Debian 12, several languages preinstalled, no network access by default (an org admin can allowlist it).
- Anthropic runs Python and Bash for Claude models on the Messages API — 5 GiB RAM and storage, one CPU, no internet, so no installing packages mid-run. Newer tool versions add state that persists between requests.
- Google runs Python only for Gemini models, with a 30-second maximum runtime and a fixed library list.
- OpenRouter runs
openrouter:shellfor any model on the Responses and Messages APIs, because the sandbox sits at the routing layer rather than inside one model provider. Both of its tools are in beta, and regional endpoints don’t offer them yet.
That last point is the interesting product decision. A routing-layer sandbox means you can swap models without touching tool code, and OpenRouter’s auto engine uses a provider’s native hosted shell where one exists and falls back to its own sandbox otherwise. I looked at the implications of provider-run sandboxes joining the API surface in an earlier post on OpenRouter’s shell tool, and this comparison is essentially the evidence for that trend.
What you give up
OpenRouter’s sandbox is billed at $0.0001 per second with a 30-second minimum for a new or sleeping container. Cheap for a handful of short commands, but the costs compound on long-running sessions.
The clearer boundary is capability. Every hosted tool here handles short, bounded work — a script, a file transform, a sanity check. Anything needing a custom base image or a GPU falls outside all four, and OpenRouter is explicit that those jobs still belong on a sandbox platform you operate. Long sessions measured in hours are in the same bucket.
Also worth noting: hosted environments change. OpenRouter’s sandbox reported Ubuntu 22.04.5 with Python 3.11.14 when they ran their test on September 22, 2026, and their own advice is to read versions from the container rather than hard-code them.
The practical takeaway
If your agent’s code execution is a few seconds of Python per request, a hosted tool removes a whole category of ops work for one JSON entry — and OpenRouter’s version keeps that flexibility across model changes. If you need GPUs, custom images, or persistent long sessions, the trade doesn’t work yet, and building on E2B, Modal, or your own containers remains the honest answer.
