AGENTIC COMMONSAI industry briefings

繁中EN

TOPICAI Chips & ComputePUBLISHED 2026-10-05

All English articlesOpenRouter

Provider-Run Sandboxes Are Becoming Part of the API: How OpenRouter's Shell Tool Fits In

If your agent needs to run code, you’ve traditionally owned the whole sandbox problem: provision a container, lock it down, keep it warm, pay for it even when the model barely uses it. On October 5, 2026, OpenRouter published a comparison of server-side code execution tools that maps how much of that burden the big providers now take off your hands.

What “server-side” execution actually changes

A server-side code execution tool runs the model’s commands inside the provider’s sandbox during your API request. The practical shift is that you stop provisioning or securing a container yourself. OpenAI, Anthropic, and Google each run code for their own models. OpenRouter’s openrouter:shell tool runs commands for any model on the Responses and Messages APIs, and openrouter:bash does the same on the Messages API.

That last part matters for how you design your tooling. If execution lives behind the gateway instead of behind a specific vendor, swapping models no longer means rewriting your execution plumbing.

Four tools, four comparison axes

The OpenRouter article compares the options on runtime, isolation, persistence, and cost. Those are the right axes if you’re evaluating this category, because they expose the tradeoffs that marketing pages tend to blur: how long your command can run before it’s killed, how firmly it’s separated from other tenants, whether state survives between calls, and what you pay per execution.

The supplied material summarizes the article rather than reproducing its full findings, so the specific numbers per provider aren’t in what I have to work from — read the original comparison for that. What the summary does confirm is that OpenRouter includes a complete request against its own sandbox along with the actual output it returned, which is more useful than a feature table when you’re deciding whether to trust a new tool.

Where you still need your own sandbox

The article doesn’t just promote its tool — it also describes the jobs that still call for a sandbox platform of your own. That distinction is worth taking seriously. Provider-run execution is convenient, but it’s scoped to a single API request, so anything that needs longer-lived state, custom dependencies, or compliance controls on the execution environment probably stays on your side of the boundary.

For builders, the decision looks like this: if your agent’s code execution is ephemeral and small — a quick calculation, a data transform, a plot — running it server-side at the provider removes an entire category of infrastructure. If execution is core to your product, with stateful sessions or strict isolation requirements, the provider sandbox is a convenience layer, not a substitute.

How this fits the routing picture

I’ve covered OpenRouter’s gateway approach before, in cheap-first routing for support bots, where the split is what your app owns versus what the gateway handles. Server-side execution extends that same split deeper into the stack: the gateway now owns not just model selection but the execution environment the model calls into.

The trend across providers is that code execution is quietly becoming a standard part of the API surface rather than a separate product you wire up. That’s good for shipping speed and less good if you assume provider sandboxes behave identically — the runtime, persistence, and cost differences between the four tools are exactly why a side-by-side comparison like this one earns its place in your evaluation notes.

If you’re building an agent today, the practical next step is to run the same workload against each provider’s execution tool and compare real outputs and latency, not just the spec sheet. The OpenRouter article shows one complete worked example to start from; your own requests are the rest of the benchmark.

Sources

AGENTIC COMMONSOperated by PHLEGON LABS
SHAREXEMAIL
Support us

Related reading

  1. Letting the API Run the Code: What Hosted Sandboxes Actually Trade Away

    OpenRouter compares four server-side code execution tools, showing where provider-run sandboxes fit and where you still need your own container.

    Code Execution

  2. Cohere's North 2 Bets Enterprises Want Agents They Can Govern, Not Just Run

    North 2 adds a redesigned agent harness, org-wide reusable skills, and admin-level spend and autonomy controls for enterprise deployment.

    Enterprise AI

  3. Letting the API Run Your Agent's Code: What Changes When You Do

    OpenRouter compares four hosted code execution tools, showing what builders give up and gain when the provider runs the sandbox.

    Code Execution