❝

IIt's Tuesday, October 6th: Welcome to another edition of The Byte.

In this piece, Maria Gorskikh takes on a question one founder asked her point blank: why would an AI agent ever need its own machine? It already runs in the cloud. The obvious answer is isolation and security. The more interesting one is that most teams never actually decide where their agents run, and they find out it mattered during an incident.

After asking dozens of founders about their biggest infrastructure problem, she found their answers sorted into three groups. Some run deterministic pipelines that belong in a shared backend. Some spin up a throwaway sandbox for each task. And a third group builds agents that drive real browsers through insurance portals, lending tools, and legacy healthcare systems, holding a customer's logged-in session and picking up tomorrow where they stopped today. Those agents need something closer to a machine of their own.

That matters because the industry tends to treat an agent as an LLM call wrapped in a loop, as if where it executes is a detail. Gorskikh argues that the execution environment decides what an agent can safely do, and that the line should follow whose data the agent touches, not how many customers you have. An agent that quietly moves from calling your own APIs to holding someone's credentials has changed categories, whether or not anyone noticed.

None of which means every agent needs its own VM. She's the first to say most don't. But it might mean the most expensive infrastructure decision is the one nobody makes.

A VM in 125 milliseconds: when your AI agent needs its own machine

A founder asked me recently why their AI agent would ever need its own VM. It already runs in the cloud. Why give it a machine?

It is a fair question. Most teams think of an agent as an LLM call wrapped in a loop. But every agent executes somewhere: in a process, on a host, next to other workloads, and often next to other customers' agents. That execution environment decides what the agent can safely do.

For most agents, the honest answer to that founder is that they do not need their own VM. But the ones that do tend to find out during an incident. The mistake we see startups make is rarely choosing the wrong runtime. It is never choosing at all.

Three ways to run an agent

Figure 1. Three ways to run an AI agent, and what each one gives you.

Shared backend. The agent runs as a worker or request handler inside your application servers, next to every other agent. It shares CPU, memory, the filesystem, and whatever credentials those servers hold. For an agent that answers support questions or calls your own internal APIs, this is the right default: cheap, simple, and easy to operate.

Ephemeral sandbox. Each task gets a fresh, isolated container or microVM that is destroyed when the task ends. This fits short jobs, especially ones that execute code nobody has reviewed. Whatever the code does, the environment is thrown away afterward.

Persistent VM per agent. Each agent gets its own long-lived virtual machine with its own disk, processes, and network identity. Files, browser profiles, cookies, and installed packages survive between runs. This is what an agent needs when it works for one customer over days or weeks.

The rule of thumb, up front:

  • If your agent follows a fixed workflow and only calls tools you built, run it in your shared backend.

  • If it executes untrusted code for a short job, give it an ephemeral sandbox.

  • If it holds a customer's sessions or credentials, or needs tomorrow what it produced today, give it a persistent VM.

Decide by what the agent executes and whose data it touches, not by how many customers you have.

What founders told us

Over the past few weeks we asked dozens of founders building AI agents the same question: what is your biggest infrastructure problem right now? Their answers mapped cleanly onto those three runtimes.

Some said they did not have one. One team runs everything on a standard workflow orchestrator with no special setup, because their agents are deterministic pipelines. Another told us their stack was sorted and that we should go talk to bigger companies. They are right. Their agents belong in a shared backend.

Others already spin up a sandbox per task and tear it down afterward. Their agents are stateless between jobs, so ephemeral environments work fine.

The third group is where it gets interesting. These founders build agents that operate software with no usable API: insurance portals where clinics check claims, lending tools where loan officers look up rates, legacy healthcare systems. The agent has to drive a real browser the way a person would. That means holding an authenticated session, keeping downloaded files on disk, and resuming tomorrow where it stopped. These agents need persistent VMs.

The founder whose question opened this piece is in that third group, and the confusion is common. "The cloud" is just someone else's computer. The real question is whether your agent shares that computer with other tenants or gets its own.

Why a shared backend is the right place to start

Three workloads turn a healthy shared backend into a liability.

Figure 2. In a shared backend, one agent that runs out of memory takes the others down with it. With a microVM per agent, the failure stays inside one VM.

First, agents that write and execute their own code. That code is model-generated and unreviewed. One customer's agent needs a different Python version or library than another's. One allocates all available memory, and the out-of-memory killer takes down the worker process along with every other agent running in it. Even without a crash, one agent running a headless browser at full CPU adds latency to every request that shares its host.

Second, agents that install dependencies at runtime. An agent that runs pip install or apt-get to finish its task is mutating a host that every other agent depends on.

Third, agents that hold authenticated sessions for customers. A session cookie or OAuth token is, quite literally, that customer's key. Keeping many customers' keys in one shared environment means building and maintaining tenant isolation in application code, forever.

There is also a human in the loop. Many of these systems enforce two-factor authentication. A person has to take over the session, approve the login, and hand control back to the agent, on the same machine, with the same browser state. That is awkward on a shared host and natural on a dedicated VM.

None of this is impossible in a shared backend. But containing it is a project, and most teams start that project only after an incident.

What a persistent VM gets you, and what it costs

A dedicated VM gives each agent a hard isolation boundary. A container shares the host kernel with its neighbors, while a microVM runs its own kernel behind hardware virtualization. A broken dependency, a runaway process, or a leaked credential stays inside one VM.

Persistence is the other half. The agent's disk holds its browser profile, downloaded files, and working directory, so a job that spans days does not have to rebuild its context or log in again every morning.

This used to be too expensive to do per agent. Not anymore. Firecracker, the microVM technology behind AWS Lambda, starts a VM in as little as 125 milliseconds with less than 5 MiB of memory overhead. On Maritime, an idle agent's VM is snapshotted and put to sleep, then restored in about a second when a message or trigger arrives, with its disk exactly as it left it.

Persistent VMs still have costs worth being honest about:

  • Operations. Someone has to build images, patch them, and debug them when they break.

  • Storage. A sleeping VM uses almost no compute, but its disk and snapshots still take up storage, and a restore is fast but not instant.

  • Shared hardware. VMs still run on shared physical hosts. They contain failures; they do not remove the need for solid infrastructure underneath.

Choose the runtime on purpose

"One VM per agent" is a good slogan and a bad rule. Your isolation boundary should follow the data, not the agent count.

Several agents working for the same customer can share one VM: same files, same sessions, same trust boundary. Agents handling different customers' data usually should not. And a startup with ten customers whose agents execute generated code may need stronger isolation than one with ten thousand customers whose agents only query a pricing API. Microsoft's multitenancy guidance describes the same fundamental tradeoff: shared deployments improve efficiency, while dedicated infrastructure increases tenant isolation and reduces noisy-neighbor risk.

In practice, most mature teams end up running all three. Orchestration lives in the shared backend, code execution happens in sandboxes, and the agents that act on a customer's behalf get persistent VMs. The runtimes are not competitors. They are layers.

Migrating later is harder than it sounds. By then, agents have left state all over the shared backend: temp files, cached sessions, packages someone installed to get a demo working. Untangling that into per-tenant VMs is a migration, and migrations are exactly what a busy startup postpones until a customer forces the issue.

Here is the quick test we use with founders. For each type of agent you run, ask three questions:

Figure 3. The three-question test for choosing an agent's runtime.

  1. Does it hold credentials or sessions that belong to one specific customer?

  2. Does its state, such as files, a browser profile, or downloads, need to survive between runs?

  3. Does it execute code nobody on your team wrote?

If either of the first two is yes, give it a persistent VM, at least one per customer. If only the third is yes, an ephemeral sandbox is enough. If all three are no, your shared backend is the right place.

The startups that get into trouble are rarely the ones who answered these questions wrong. They are the ones who never asked, while their agents quietly moved from the first category to the third.

Why we built Maritime

At MIT, my research involved benchmarking AI agents, which meant running many isolated agents in parallel on AWS. It was slow to set up and expensive to run. That experience is why we built Maritime.

Maritime gives every agent its own isolated microVM with a persistent disk, so its files and sessions are still there on the next run. Agents sleep when idle and wake in about a second, which is what makes one VM per agent affordable.

Start with a shared backend. Then give your agents their own VMs on purpose, before a customer's incident decides for you.

The AI Collective is built by volunteers across 180+ chapters in 40 countries.

Thank you to the thousands of volunteers around the world who make this work possible. We truly could not do this without you.

🧑‍💻 About the Author & the Editorial Team

Maria Gorskikh is the co-founder and CEO of Maritime (YC F26), which gives every AI agent its own isolated, persistent computer. Before Maritime, she researched AI agents at MIT, including agent communication protocols and agent benchmarking.

About Josh Evans

Josh is a Managing Editor at The AI Collective Newsletter and leads content for The Byte. Outside of AIC, Josh works in Content Protection at Spotify.

Add Your Thoughts

Avatar

or to participate

Keep Reading

View more