Worker status

What each worker status means, who writes it, and exactly what happens when a container dies or a node disappears.

Every worker carries one status. It is the short answer to "what is this machine doing right now", and it is what the dashboard pill, the command palette and the API all read.

What makes it worth a page of its own is that the status has several authors. You write it when you click a button. The platform writes it while it prepares the machine. The node your worker runs on writes it through its agent, which is the only thing close enough to the container to see what actually happens to it. Every diagram below names, next to each status, who put the worker in it.

You create a workspace, it comes up, and much later you delete it. That is the whole story, and five statuses tell it.

YOU CLICK NEW WORKSPACEprovisioningnothing has started yetthe platformno image for this project version yetbuildingyour project's image is bakedthe platformthe image is readypullingthe image reaches the machinethe nodethe container has been createdstartingsetup runs inside the containerthe nodesetup finished — code-server answersreadyyours to usethe node… LATER, YOU CLICK DELETEdeletingthe node tears it downyoucontainer and volumes removed— gone —nothing left to showthe platform
The nominal path. Each status carries who put the worker in it.
StatusWhat it means
provisioningThe worker exists as a row. Nothing has been created yet.
buildingYour project's image is being baked on the node. First worker on a given project version only — the others reuse the image. This is the one that can take minutes, and the Logs tab shows the build log while it lasts.
pullingThe image is landing on the machine that will run your container.
startingThe container is up and the setup is running inside it: credentials, clones, postCreateCommand. The progress bar follows it phase by phase.
readyYours. code-server, ports, terminals and SSH all answer.
deletingYou asked for it to go. The node is removing the container and its volumes.

Why so many steps before starting?

Because each one can be the slow one, and you deserve to know which. A ten-minute image build and a ten-second container start both used to look like "setting up" — now building says which one you are waiting on, and gives you a log to read while you wait.

Stopping is not deleting. The container is stopped, the volume stays, and everything you left in /workspace is there when you come back.

STOP AND START ARE A ROUND TRIPstoppingyou click Stopstartingyou click Start — postStart runs againreadycontainer runningstoppedvolume kept, disk intact
Stopping does not destroy anything — the volume and everything in /workspace stay put.

A restarted worker runs postStartCommand again but not postCreateCommand — so it stays in starting until your start-up command finishes, however long that takes.

Four different bad things can happen to a worker, and they are four different statuses. The point of keeping them apart is that they don't call for the same reaction: one asks you to restart, one asks you to fix your postCreate, one asks you to look at a build log.

THE CONTAINER DIES ON ITS OWNreadyout of memory,or it just exitsexitedexit 137 · OOMstart it againSETUP FAILSstartingpostCreate returnsa non-zero codeerrorsetup-failedphase + message keptTHE IMAGE BUILD FAILSbuildinga feature failsto installerrorbuild-failedbuild log still thereSOMEONE DELETES THE CONTAINERreadymissing from thenode's inventoryerrorremoved-externallycaught within 30 s
Four bad outcomes, four distinct statuses — and in each one, something that says why.
  • The container dies on its own — out of memory is the usual reason. A dev container that compiles something big, or a test suite that eats the box, gets killed by the kernel. Nothing was wrong with your setup: start it again, and give the node more room if it keeps happening.
  • Setup fails — a clone that can't authenticate, a postCreateCommand returning non-zero. The phase it died in and the error are kept on the worker, and the Logs tab has the output.
  • The image build fails — a devcontainer feature that won't install, a base image that doesn't exist. The build log stays available on the project, so the failure is readable after the fact.
  • The container is deleted on the node — someone ran docker rm on the machine, or a cleanup script did. The platform notices at the next reconciliation and says so rather than pretending the worker is still there.

Your workers run on your compute. A node can lose its network, be powered off, or simply have its agent restarted during an upgrade. None of that is an error, and none of it destroys anything — but for as long as it lasts, nobody can see what your containers are doing.

THE NODE DISCONNECTSnetwork drop, machine powered off, agent restarted…provisioningno container yetthe wait died withthe connectionerrornode-loststart a new onereadycontainer existsafter a graceperiod (~60 s)unknownwe no longer knowit remembers what it wasTHE NODE IS BACKreadystill runningexiteddied meanwhileerrorno longer therethe node's inventory decides — it says what actually exists on the machine
A node going away breaks nothing — it makes the status uncertain, and the reconnection settles it.

Two very different things happen, depending on whether the container existed yet:

  • It did not — the worker was still being prepared, and that preparation lived in the platform's memory. It died with the connection and nobody will pick it up. The worker goes to error, and the honest fix is to start a new one.
  • It did — the container is almost certainly still running on the machine; we just cannot see it. The worker becomes unknown, which claims nothing, and remembers what it was.

When the node reconnects, its agent sends a full inventory of what is actually on the machine, and that inventory decides: still running means ready, died in the meantime means exited, gone means error. Nothing needs you to intervene.

StatusMeansWritten byWhat you can do
provisioningRow created, nothing startedThe platformWait
buildingProject image being built on the nodeThe platformRead the build log
pullingImage landing on the machineThe nodeWait
startingContainer up, setup runningThe nodeFollow the progress, read the logs
readyUsableThe nodeEverything
stoppingStop requestedYouWait
stoppedContainer stopped, volume keptThe nodeStart, rebuild, delete
deletingBeing removedYouWait
exitedContainer died on its own — statusMeta says the exit code, and whether it ran out of memoryThe nodeStart it again
unknownNode unreachable, status uncertain — statusMeta.previousStatus keeps what it wasThe platformWait for the node
errorAn operation failedThe platformRead the reason, then rebuild or delete

Every change of status is also written to a journal, in the same step as the status itself, so the two never disagree. The History section of a worker's page reads it: a strip where each segment is as wide as the time spent in that status, and below it the list of transitions, newest first — what the worker became, for how long, who wrote it (you, the node, the platform), and why when there is a why ("Ran out of memory (exit 137)").

It is the answer to "why did this worker die on Tuesday" once the container that could have told you is gone. The journal goes away with the worker.

  • How far the setup got. That is a separate, finer-grained thing: while a worker is starting, a per-phase status says whether it is cloning, running postCreate, or starting code-server. The status says what the platform is doing; the progress bar says how far the setup inside the container has got. During an image build the second one has not started yet — which is why you can see "Building image…" above a bar that has not moved.
  • Whether the node is healthy. Reachability belongs to the node, not to your worker: the Nodes page is where you see if a machine is online, and since when.

Tip

Everything on this page is visible through the API too — GET /api/orgs/{orgId}/projects/{projectId}/workers returns the status of every worker in a project, GET /api/orgs/{orgId}/projects/{projectId}/workers/{workerId}/status-history its journal, and the MCP server exposes the same thing to an agent.