When you build AI features that handle internal company data, a requirement sometimes comes up along the lines of: "The aggregation conditions change every time, so we want it to write and run code on the spot."
This is where things change. If you only generate text, calling an external API and returning the result was enough. Once you run generated code, you need a place to run it. What is more, the code being run was not written by a person.
What providing a place to run code actually means
What this kind of feature needs is an execution environment that meets all of the following at the same time.
- Code being executed cannot reach other users' data or other internal systems
- Code starts running shortly after a request arrives
- No costs are incurred while it is not in use
- An infinite loop or a large memory allocation does not drag down other processes
It may look as if spinning up a single container would do, but trying to meet all four of these turns it into platform design. We covered the cost of idle time in Agents spend a long time waiting. This time we look at the step before that: where to draw the line on whether to build it yourself at all.
Looking at one example taken to the extreme
Colin Weld and Connor Adams, engineers at Modal, which provides sandboxes among other things, published an account of rebuilding the company's sandbox platform on its blog on July 16, 2026. InfoQ reported on it on September 23.
According to the account, the company was already running millions of sandboxes a day and supported up to 50,000 concurrent sandboxes per customer. Reinforcement learning, on the other hand, can require running millions of sandboxes at once and creating hundreds of thousands of them together at the start. The starting point for the rebuild was that the existing platform had not been built with this scale in mind, and that the same was true of other existing systems.
The company explains why conventional container platforms get stuck at this scale, using Kubernetes as an example. In the worst case, scheduling grows with the product of the number of nodes and the number of pods, and by default it is processed one at a time. Each pod causes multiple writes to etcd, the central durable store, between its creation and its removal, and etcd cannot natively be sharded within a single keyspace. Just signaling that nodes are alive keeps up a stream of writes proportional to the number of nodes. Scaling up is possible, the explanation goes, but it takes serious work such as rewriting or replacing etcd and parallelizing scheduling.
Modal itself does not build on Kubernetes, but its original platform had similar problems. Because it relies on strong consistency across the whole backend, every time a sandbox was created and placed, global coordination was needed, and writes proportional to the number of sandboxes piled up in a Postgres instance that could not easily be sharded, the company writes.
The direction of the rebuild was to stop this central coordination. On the path for creating and running sandboxes, it gives up global consistency: a fleet of scheduling servers running in parallel picks a worker to place each sandbox on from state held in memory, and asks that worker directly over RPC to create it. The worker accepts if it has free capacity and rejects otherwise. Each worker holds the source of truth for its own state and periodically publishes it to a Redis stream. No data store sits in the creation path, and records are mostly written asynchronously afterward. That is the architecture.
In the company's measurements, 1 million sandboxes could be created in under a minute (it says the bottleneck was on the benchmark side), and the time from requesting creation until the user's code could run was under 0.5 seconds at the median. The slow tail was longer than expected, and the company attributes much of it to kernel and network contention when many sandboxes start at the same time on the same worker. The most imminent bottleneck, it explains, is the single Redis stream to which all workers publish their state, and load testing suggests it will hold up until well over 100,000 workers.
The new platform was in beta as of the July post. According to the release notes for the company's Python SDK, it can be selected with the environment variable MODAL_SANDBOX_V2=1 starting with 1.5.4 on August 12, and it is scheduled to become the default in 1.6.0 (the latest version as of September 28 is 1.5.5).
All of these figures are the provider's own measurements, and our editorial team has not reproduced them.
The takeaway is not the number
The value to take away from this account is not "1 million." It is the design property that in a design that keeps a single source of truth at the center, processing proportional to the number of sandboxes and nodes concentrates there, and as scale grows, that point becomes the bottleneck.
And what the company spent to remove that bottleneck was months of work spanning most of the major systems in its backend. Besides the fleet of scheduling servers, worker-side state management and the RPC path, it included rebuilding every sandbox feature and all monitoring, and changes to worker management and the container runtime. When starting large numbers of containers at once, it also ran into a problem where they contended for a Linux kernel lock during network setup and took tens of seconds to start, and it changed the network configuration for sandboxes. This is the work of building a product, an execution platform, not the work of building a feature attached to a client's business system.
On the other hand, suppose an internal feature peaks at tens of concurrent runs. The official Kubernetes documentation assumes up to 5,000 nodes and 150,000 pods per cluster, and Modal's platform before the rebuild already supported up to 50,000 concurrent sandboxes per customer. Tens of runs is several hundred times smaller than either figure, or more. This is not a scale at which the central-coordination bottleneck covered in this account becomes the core of the decision. At this scale, what tends to run short first is the people available to look after the platform. That is our editorial team's assessment.
Why you might want to build it yourself, and what those reasons really are
Even so, when a proposal to build it in-house comes up, the following three reasons are conceivable. All of them sound right at first.
"We can't send client data to an external service." This may be a real constraint, but it is not a reason to build the execution environment itself. Using a managed execution environment that runs inside a cloud account or project your company manages may satisfy the boundary requirements. We covered Google Cloud's Cloud Run sandboxes (in preview as of September 28) in Is AI-generated code safe to run directly in production?, and AWS Lambda MicroVMs (the June announcement also covered the Tokyo Region) in Overcoming stateless Lambda limits. What you can choose depends on whether what is prohibited is "handing it to a provider" or "taking it out of the country."
"If startup is slow, people won't use it." Modal writes that the main reason startup got faster on its new platform is that scheduling now takes only tens of milliseconds. The company is still working to shorten the slow tail when many sandboxes start at once. If you build your own because of startup speed, you will also own the work needed to maintain that speed (the scheduling path, container startup and network configuration). What you should decide first is how many seconds users can wait.
"We want to prepare for future scale." If what you are preparing for is concurrency, first estimate whether that scale will really arrive. Modal itself judged, once the scale being demanded had changed, that rebuilding from scratch would be faster than growing its existing platform. If you build ahead of time, all that remains is operating for a scale that never came.

The vertical axis of the diagram is whether you can restrict what runs, and the horizontal axis is concurrency. Start from the top left (processing can be restricted, concurrency up to the double digits), and only when you reach the bottom right (arbitrary code, concurrency in the triple digits or more) do you start comparing platforms. The digit thresholds are our editorial team's rule of thumb.
Where to draw the line
When building an execution environment into an AI feature in client development, checking the following in this order makes the decision easier (this is our editorial team's framework, not an official recommendation).
- Put a number on how many runs you expect at the same time. If peak concurrency stays in the double digits, we think there is little need to build your own platform because of concurrency. First, check whether a managed sandbox or a job execution service is enough.
- Agree with the client, in words, on how strong the isolation needs to be. The options change depending on whether "it is enough that other companies' data is not visible" or "we want separation down to the host kernel." We explained in Not relying on VM isolation alone that isolation by VMs also rests on assumptions.
- Translate startup speed into a user-experience requirement. A "median of 0.5 seconds" is an attractive number, but if a few seconds on the screen where users wait for results is no problem for the feature, paying for that difference is worth less.
- First consider whether you can restrict what can be run. If you can move away from arbitrary code toward combinations of predefined operations, the requirements for the execution environment drop a level. If you design on the assumption that "anything can be run" without working this out, only the isolation requirements keep piling up later.
- Decide who will operate it first. Adding one more sandbox platform adds one more thing to monitor, handle incidents for and update. If it has not been decided who will look after it after handover, it is safer to lean toward not building it. We have also covered managed options in our articles on GKE Agent Sandbox and Cloudflare Sandboxes.
What to do next
For the feature you are considering now, write down peak concurrency as a single number. If it stays in the double digits, we recommend spending your time on discussions to narrow the range of what can be run rather than on designing an execution platform.
If it is expected to reach the triple digits or more, and to stay there, only then does comparing managed products with building your own start to make sense. The axis to compare on is not performance but "who will operate it."
On September 28, 2026, we directly checked Modal's engineering blog post (dated July 16, 2026), the release notes for its Python SDK, the official Kubernetes documentation "Considerations for large clusters" (v1.37), the Cloud Run sandboxes documentation and the Lambda MicroVMs announcement. We referred to InfoQ's article dated September 23 as a news report. The figures of 1 million sandboxes, under a minute, a median under 0.5 seconds and 100,000 workers are all the provider's own measurements, and our editorial team has not reproduced them. We have not run any managed execution environment, including Modal, and have not compared startup speed, cost or isolation strength. The split by the number of digits of concurrency and the five decision points are our editorial team's framework.
For choosing an execution environment for AI features or building one into existing systems, please consult GleamHub.
Sources
- Scaling to 1 million concurrent sandboxes in seconds — Modal Blog
- Python SDK Release Notes — Modal Docs
- Considerations for large clusters — Kubernetes Documentation
- Code execution in Cloud Run(Cloud Run sandboxes) — Google Cloud Documentation
- AWS introduces Lambda MicroVMs for isolated execution of user and AI-generated code — AWS What’s New
- News report: Beyond Kubernetes at Modal: How to Scale 1 Million Concurrent Sandboxes in Seconds — InfoQ








