Mainbrella uses idempotency keys to stop agent retries spawning duplicate machines

Founder Andrew Arrow's October 7th engineering post details how the service handles lost replies, delayed commands and other failures in cloud Linux workspaces.

By · Published

Primary source: Mainbrella Engineering

Why it matters

Agent sandboxes must recover from lost replies without duplicating machines, commands or compute charges. Mainbrella's open-source implementation makes retry rules inspectable, while its published test checks exclusive admission and parallel provisioning, not reliability at production scale.

Mainbrella uses idempotency keys to stop agent retries spawning duplicate machines — Founder Andrew Arrow's October 7th engineering post details how the service handles lost replies, delayed commands and other failures in cloud Linux…

Andrew Arrow is building Mainbrella around a small but expensive failure: a cloud machine starts, the response gets lost, and an agent retries by creating another one. In an engineering post published October 7th, Arrow lays out the controls Mainbrella uses to make those retries safer, from reserving capacity before boot to refusing to replay work when the system cannot tell what happened.

Mainbrella sells Linux workspaces that AI agents can create and control through an API. Its homepage lists its Builder plan at $5 per month. The post argues that making a machine available is one part of the service: an agent also needs to know which machine it owns, whether a command ran, and whether retrying will duplicate work or spending.

Arrow's GitHub profile and LinkedIn profile describe a long career in software engineering and list the University of Pittsburgh alongside work connected with Yammer, Bird and Shopzilla. He operates Mainbrella as a sole proprietorship in Culver City, California, according to the company's homepage. The post focuses on backend work that determines whether an automated job can be trusted after a network failure.

Reserve first, boot second

Mainbrella's open-source backend uses a Cloudflare Durable Object as an account-level coordinator. Before a machine is dispatched to boot, the coordinator reserves a slot, a monthly start and compute allowance. Separate objects handle individual machine slots. That ordering prevents two simultaneous requests from both seeing the same last available slot and starting machines before either request updates the account's capacity.

Diagram of Mainbrella's account-level coordinator reserving capacity before releasing the admission lock and allowing a machine to boot, with separate objects handling individual machine slots.
Mainbrella's documented admission flow reserves capacity before boot; separate objects handle individual machine slots - AI explanatory diagram, not documentary evidence. RuntimeWire - AI-generated diagram.

The reservation is deliberately brief. Once it is durably recorded, Mainbrella releases the admission lock and lets the runtime boot independently. A test in the repository sends six simultaneous creation requests against a Builder account with five slots. It holds readiness checks behind a gate; five machines reach boot, while the sixth request gets a conflict. The test demonstrates admission control and parallel provisioning; it does not measure production startup speed.

The client also sends an Idempotency-Key with a creation request. Mainbrella saves a receipt for that attempt alongside its reservation in one atomic storage write. If the reply goes missing, a retry using the same key can find the original attempt rather than claim a new slot. When the reservation is still pending, the coordinator waits through a 90-second reconciliation window, then checks what exists at the runtime. If the machine is running, it returns that machine without charging another start.

Diagram showing how Mainbrella saves a creation receipt with its reservation, uses the same idempotency key to find the original attempt after a lost reply, and checks the runtime during a 90-second reconciliation window.
The post describes receipt-based retries and a runtime check; a running machine is returned without another start charge - AI explanatory diagram, not documentary evidence. RuntimeWire - AI-generated diagram.

The retained key lasts 24 hours. If the machine has stopped or its slot has been reused, Mainbrella returns creation_no_longer_running; the client must use a new key to request a replacement. An ambiguous retry does not silently become a second machine.

Delayed commands need an identity too

A request can arrive late enough to target a slot that has already been reassigned. Mainbrella gives each admitted start an increasing reservation number and records cancellation fences at the runtime. A delayed boot for an older reservation is rejected, and a late cleanup for that old machine cannot stop its replacement. Commands and file requests also include a machine generation, so a reusable slot name alone cannot authorize an operation against a different machine lifetime.

The same caution shapes command execution. Mainbrella's post describes managed jobs with their own idempotency keys and saved execution records. Clients can reconnect to read output events from a stored sequence cursor after a dropped connection. If the runtime restarts before a job finishes, Mainbrella marks it interrupted and does not rerun the command automatically. For a script that writes a CSV, that can mean losing unfinished work. For an operation with side effects, an automatic replay could be worse.

A readiness check runs uname -a and confirms that the guest can execute a command; it does not prove that an agent's application ran correctly. Likewise, a completed file transfer does not prove that a calculation used the right inputs. Mainbrella puts lifecycle and identity controls around a workload, while leaving application correctness to the workload itself.

The promise stops at the durable boundary

The post follows the same logic into saved workspaces. Capturing a disk involves the account coordinator, the runtime and the cloud provider, which cannot commit each other's state in one transaction. Mainbrella records an operation before requesting a capture and saves the provider's handle before telling the account that the workspace is ready. If the runtime loses its reply after saving the handle, a retry can recover the same capture.

There is a harder failure: the provider may accept a capture, then the runtime may stop before saving the handle needed to restore it. In that case, Mainbrella blocks another capture for the unresolved operation and leaves the original machine running. The snapshot may exist, but without the handle the service cannot safely promise it can restore the snapshot.

The post's examples and tests demonstrate control flow rather than measured live performance. For a young service competing for agent workloads, the post makes reliability behavior inspectable in code. The public evidence establishes what the implementation is designed to do, not how often it has been tested under real customer load.

Reader comments

Conversation for this story loads after sign-in.