Skip to content
Putting technology to work.
Insights to guide decisions and action.

Search articles

Agents spend a long time waiting — the assumption AX resets

Table of contents · 7 items

In a setup where an internal AI agent runs in a container that is always on, you pay for the runtime environment for the whole time it is running, even if the time spent actually processing is short.

The cause lies in the shape of the work. An agent's work is interspersed with time spent waiting for model responses, results from external APIs and human approval. Google, too, explains that agents often run in short, intensive bursts followed by long waits, and that keeping them running during those waits wastes compute. The CPU is barely used while the agent waits, but in an always-on setup the container stays allocated.

A different starting assumption

On May 20, 2026, Google announced "Agent Executor," an open-source runtime for agent execution, resumption and distributed deployment, on the Google Cloud blog. It is published as google/ax under the google organization on GitHub, and the command name is also ax (hereafter, AX). It is licensed under Apache 2.0, and the blog says it is available as a preview. InfoQ also reported on it on September 22.

AX is not a framework for building agents. It is a runtime for running the agents you have built in sandboxes, preparing their working environment and running them in larger numbers. Google's blog says it is not tied to any particular agent framework (harness). Running it requires a Kubernetes cluster with Agent Substrate installed; Google released Agent Substrate on the same day.

The README positions agents as "a new kind of workload, neither stateless microservices nor run-to-completion batch jobs." The official site goes into more detail. Agents are stateful, bursty and long-running. They compute intensively for a while, then wait for responses from models and tools or for human approval. According to the site, with traditional orchestrators, keeping sandboxes running while they wait drives up costs, and there is no built-in way to suspend and resume them in under a second.

In other words, it puts "how to keep the waiting state cheap" at the center of its design.

How suspend and resume actually work

The README lists three basic building blocks: Task, which runs work in a sandbox with CPU and memory limits; Workspace, which prepares Git repositories, MCP servers and skills in advance; and Model, which brings together the model to use, its settings and its credentials. The official site lists four, adding Gateway, which restricts outbound destinations to an allowlist, but as of September 26 the repository's documentation (README, Concepts, Manifests, API reference) did not describe Gateway.

What matters for waiting is how tasks are suspended and resumed. Based on the repository's documentation, the current behavior is as follows.

  • ax suspend saves the state and suspends a task, and ax resume resumes it. The README describes this as a way to pause an idle agent and resume it from where it stopped.
  • When a request arrives for a suspended task, the Agent Substrate router resumes the task first and then forwards the request.
  • However, what remains after a suspend are the files in the working area (/workspace). On resume, the process starts again in a new container, so any state needed to continue must be written to files before the task stops.
  • Automatically detecting idle time and suspending tasks is still listed on the roadmap as a future item.

The official site says idle agents can be suspended and brought back in under a second. We have not run AX and have not verified this speed. For now, it is safer to assume that suspending is done through commands or the API.

Where it matters for custom development

From the standpoint of building agents for your own company or for clients, the following points matter.

Idle cost becomes something you can design for. To the two options of keeping an agent always on or starting it each time and accepting the slow startup, it adds a third: suspending it and bringing it back when needed. The prerequisite is that the agent is built to continue where it left off after a resume.

Processes that include human approval become easier to build. Waiting for approval can drag on. Google's blog also explains that long-running execution needs to be able to resume after interruptions caused by outages or human confirmation. If an agent does not have to hold on to resources while it waits, agents that include an approval flow become easier to build. We covered where to place approvals in our article on putting approvals in code.

But it also adds to what you have to operate. The README's prerequisites are a Kubernetes cluster with Agent Substrate installed, Go and kubectl, ko for building container images, and a container registry the cluster can pull from. The control plane is deployed to the cluster together with Redis. Adding one more platform adds one more thing to monitor and respond to when it fails. Our view is that with only one or two agents, this operational burden is likely to outweigh the savings on idle costs.

Do not base the decision on scale figures

The README and the official site say AX is designed to run billions of tasks in a single cluster. The Agent Substrate README likewise claims 10x the density of standard container runtimes and resume times under 500 milliseconds. These are all the providers' own descriptions, and we have not measured them.

What you need for the decision is not the upper limit of scale but how much waiting time there is in your own setup. That is something you can measure yourself.

How to tell whether it is worth evaluating

Before considering adoption, checking the following makes the decision easier.

An editorial concept diagram showing how the decision splits: measure in your own environment the share of run time spent waiting (the ratio of time spent waiting on models, external APIs and human approval); if it is low, start by reviewing your runtime environment settings, and if it is high and the number of agents is expected to grow, consider a trial. Not an actual measurement result

  1. Measure the share of waiting time for the agents you run now. This is the proportion of run time spent waiting on models, external APIs or people. If it is small, you gain little from a suspend-and-resume mechanism.
  2. Count how many agents run at the same time. While the number is small, before adding a new platform, first check whether your current runtime environment offers settings that stop instances during unused time or reduce the number of instances.
  3. Look at what the waiting is for. If most of it is waiting for human approval, revisiting the design of the approval flow can do more than changing the mechanism.
  4. Check whether your agents can continue where they left off after a resume. With AX's current mechanism, the process restarts on resume and the files in the working area remain. If an agent keeps its state only in memory, it will need rework.
  5. Factor in that it is at the preview stage. The README states explicitly that major breaking changes are likely before a stable release. Agent Substrate is also pre-1.0 and does not guarantee backward compatibility. It is safer to put it into a client's production environment only when you can set aside a separate period for testing.

Designing the agent's work procedure itself (which steps to delegate and where to hand back to a person) is a separate layer from the runtime. We cover the design of setups that combine multiple agents in our article on designing agent orchestration.

What to do next

If you have agents running, calculate the waiting-time ratio from a week's worth of logs. That number is the first input for telling whether this topic is relevant to your company.

If the ratio is high and you expect the number of agents to grow, it is worth trying it out in a test environment. If not, reviewing your runtime environment's settings comes first for now.

We checked two Google Cloud blog posts (the Agent Executor and Agent Substrate announcements, both dated May 20, 2026), the official AX site (agentexecutor.io), the README, documentation (Concepts, Manifests, Runners, Networking, Roadmap, Architecture) and license of google/ax on GitHub, and the Agent Substrate README and Architecture documentation on September 26, 2026. We have not installed or run AX, measured how fast it suspends and resumes, or compared costs. The providers' figures on scale and resume speed, and the Gateway specification that appears only on the official site, could not be verified, so we have not used them as a basis for decisions. InfoQ's article dated September 22, 2026 was used as a news report.

For designing the architecture of internal AI agents or estimating costs including idle time, please consult GleamHub.

Sources

Share this articleXFacebook
Kakeru Suzuki

Fascinated by the possibilities of technology, has had a deep interest in programming and digital art since student days

Turn this article's theme into your company's next step

Concrete steps forward for your organization.

We organize your desired architecture, legacy systems, and operational requirements to formulate your next steps toward execution.

  • Desired architecture
  • Integration with existing environments
  • Operational requirements
Consult on development & operations initiatives

You can consult with us from the initial conceptual stage. Details from this article will be carried over to the inquiry form.

Receive the latest articles by email