Advanced Tasks

Running a different agent, writing your own session stream, driving tasks over the API, and the reasoning behind the parts of the model that look strange.

Everything in Tasks works without reading this page. This one is for when you want to run something other than Claude Code, drive the whole thing from a script, or understand why a task refuses to close itself.

The only contract is that your command reads the prompt on standard input:

my-agent --headless --yes

Set it under Edit → Advanced options → Lifecycle → Delegated work → Agent command. It runs inside the task's worker, in the first repository's directory, with the project's secrets already resolved.

Spunto adds exactly two things to that command, and only when asked:

  • the flags that make a harness emit a machine-readable stream (below) — skipped entirely if your command already sets --output-format itself, so a pipeline stays yours;
  • --model <id> when the task resolved a model — skipped if your command already names --model.

Agent command, Session stream and Follow-up command are picked from a row of cards — today Claude Code plus Custom — served by GET /api/harness-packs. Picking a card fills in all three fields and hides them; there's nothing to read that the card doesn't already say. Pick Custom to bring them back, pre-filled, and edit from there.

A pack is a suggestion, never a stored choice. Nothing on the project records which card you picked: the form re-derives it on every render by comparing the current fields against the catalog. So editing a command by hand is always safe — it just stops matching a preset, and the form falls back to Custom on its own.

The opening turn the harness receives is your prompt plus a two-line block, and the split between those two lines is the whole design:

LineKind of thingConfigurable
Spunto context: you are running in a disposable worker, on branch …, already checked out.a fact the agent can't infer from inside its containerno
When the work is done: commit on this branch and push it.an opinion about what "done" meansyes — Settings → Agent instructions, per organization
(nothing by default)what only this repository needs saidyes — Agent instructions in the project's task settings, added after the organization's

An org's instruction replaces the default rather than appending to it. Appending would leave two sentences contradicting each other the moment someone wants a different workflow — and the one they didn't write would be first.

The organization's instruction is a house rule — what "done" means has no reason to differ between two repositories of the same team. What does differ (which test suite to run, which folder not to touch) goes in the project's own Agent instructions, which is appended after the organization's rather than replacing it: a project can add to a house rule, not quietly opt out of it.

It's read when a task starts, never copied onto the task row. A running task keeps the instruction it began with — re-reading it every turn would silently rewrite the history of sessions already in flight.

Note

Skills are files, not text, so they don't go through this instruction: they come from the organization's skill sources, refreshed before every task. Sub-agent definitions and settings.json still reach the agent through the cloned repository or the member's dotfiles.

Every command Spunto runs for a task — the session, the follow-up, reset, accept, drop — gets these:

SPUNTO_TASK_ID              SPUNTO_TASK_BRANCH        SPUNTO_TASK_STATE
SPUNTO_TASK_TITLE           SPUNTO_TASK_BASE_BRANCH   SPUNTO_TASK_SESSION_STATUS
                                                      SPUNTO_TASK_MODEL (when resolved)

The follow-up command gets one more, SPUNTO_TASK_SESSION_ID — the harness's own session id, which is what makes resuming a conversation possible.

Your project's secrets are there too, resolved exactly as they are for a worker you open yourself — so an accept command can be gh pr merge without you pasting a token into the project settings.

Note

The SPUNTO_TASK_* variables are applied last. A secret sharing one of their names cannot override it — they are the contract your scripts rely on.

Default model on the project and Model in the New task dialog are the same dropdown, fed by the pack catalog. Neither is a closed set: the field stays free text underneath, because the valid model ids belong to whatever your agent command actually runs, not to Spunto. A project on Codex — or on a script of its own — names models this catalog has never heard of.

Whatever you pick or type is appended as --model, and also exported as SPUNTO_TASK_MODEL for a harness that reads models differently. A task's model is resolved once, at creation, and stays fixed for the whole conversation including every follow-up turn — rather than being re-read from the project on each turn.

An agent session lasts minutes to hours. Without a machine-readable stream, Spunto would have exactly one fact to offer over all that time — "it's running" — with the output arriving in one block at the end. Session stream is the setting that changes that.

ValueWhat it does
claude-stream (default)Interactive. --output-format stream-json --verbose --input-format stream-json. The session stays alive between turns, so you can read it live, talk to it while it works, and interrupt a turn without killing it.
claude-json--output-format stream-json --verbose only. The session runs once and hands back. Readable live, but every reply is a fresh resume.
jsonlYour harness already emits Spunto's vocabulary, one JSON object per line. Nothing else to learn.
noneNothing captured beyond the tail of stdout, readable once it's over. For a harness that can't emit anything.

none is the right setting for a harness that genuinely can't produce a machine-readable stream, and the wrong one for anything that can — you pay the same money for strictly less.

What gets stored is never a harness's shape:

session.started   session.title   message      thinking
tool.call         tool.result     plan         usage      session.ended

…plus raw for anything an adapter couldn't place.

There has to be a common denominator: you cannot render Codex, Gemini CLI and a shell script of your own through one vendor's message format. Adapters translate at the boundary, and raw guarantees nothing is lost silently — an unknown event is itself information about the harness.

The harness's own line is kept next to every event. Normalizing without it would be a one-way door: an adapter that maps a field wrong, or hasn't learned a new one yet, would destroy the original on write. With it, type and payload are a projection — re-derivable from the line that produced them. It's expensive relative to what it explains, so it's never returned unless you ask for it (?source=true).

Which is also why there's a third option, Spunto events (JSONL): print one of those objects per line from your own harness and it renders exactly the same way.

{"type":"message","payload":{"role":"assistant","text":"Found it."}}
{"type":"tool.call","payload":{"callId":"1","name":"Edit","input":"{\"file_path\":\"src/a.ts\"}"}}
{"type":"tool.result","payload":{"callId":"1","output":"1 replacement.","isError":false}}

Note

Print {"type":"session.title","title":"…"} and your task gets named the same way Claude Code's conversations do.

curl -H "Authorization: Bearer spk_..." \
  "https://spunto.net/api/orgs/{orgId}/projects/{projectId}/tasks/{taskId}/events?since=0"

seq is dense and monotonic per task — send the highest one you have back as since and you never re-read what you already have. Add &source=true for the harness lines.

The Session panel shows context, tokens and cost for the whole task, not for the page of events you happen to be reading. Three rules behind those numbers:

  • Tokens are summed by us, over per-call events. A turn in flight hasn't produced a final result line yet — a recap that waited for one would show nothing during exactly the period someone is watching.
  • A harness's own running total is never added up. In interactive mode a result line lands on every turn and repeats the whole session. Summing them would bill the first turn once per later turn.
  • Money comes from the harness, never from us. Pricing a call needs a price list; a platform that invented one would be showing a number it can't defend.

The context gauge is the last call's prompt against the window the session opened with. Nothing in the stream states that window, so it's resolved from a catalog of model ids, or from the Models API when the task's secrets carry an Anthropic credential. A model neither source knows gets no percentage — a gauge against a guessed window is worse than no gauge, and the token count stays correct either way.

POST .../tasks/{taskId}/messages    { prompt, files? }   → 202
POST .../tasks/{taskId}/interrupt                        → 202

On the default interactive stream, the session holds a live stdin inside the worker, so a message is written straight into it — no job, no resume, no restart. The session that already holds the context is the one that receives it. Interrupt sends a control request: the harness ends the current turn and stays alive, so the next message is an ordinary turn. That's the whole difference with Drop, which kills the process and fails the task.

On the non-interactive streams there's no open stdin, so a reply is only accepted while the task is in-review — which is when it's waiting for you anyway — and it works by resuming the harness with its session id. Spunto only knows that id because it read it off the stream, which is why answering requires a session stream at all: the alternative would be a fresh session with no memory of the turn it's meant to continue.

A harness that resumes differently gets a Follow-up command of its own.

Note

Known rough edge: an interrupted turn shows up as a failed session.ended. That's true — the turn didn't finish — but the timeline doesn't yet distinguish "interrupted" from "crashed".

Files travel both ways, all types. You attach them from the clipboard (where a screenshot lives), by dropping them anywhere on the composer, or from the paperclip at the left edge of the writing area. The agent's own files — a screenshot from a tool, a PDF it read — are pulled out of the stream and stored, with the line keeping their id in their place.

They reach the agent by path, never by payload: each file is written into $HOME/.spunto/attachments/<task>/<id>/ in the worker, and the text sent to the harness gains a line naming those paths. So a harness needs to know nothing beyond "open a file I was given the path to" — which works for every dialect, none included, for a PNG as much as for a 4 MB CSV, and the file is still there to be reopened later.

Limits: 10 MB per file, 10 files and 25 MB per turn, deduplicated per task on content, deleted with the task. Images are downscaled to 1568 px before upload — what the model resizes to anyway — so a retina screenshot leaves under 200 KB. Everything else goes as-is.

Warning

Only PNG, JPEG, WebP and GIF are served with their own content type. Everything else downloads rather than rendering, whatever type it was uploaded as — serving user-supplied HTML or SVG from the dashboard's own origin would be stored XSS against the whole organization. An SVG is accepted fine; it just downloads instead of executing.

Every command Spunto ran for the task, with its output — the branch setup, the agent session, the accept or drop command. Each shows the exact shell that ran, where, how long it took and what it printed. Failed and still-running ones are expanded.

curl -H "Authorization: Bearer spk_..." \
  https://spunto.net/api/orgs/{orgId}/projects/{projectId}/tasks/{taskId}/commands

Oldest first — the order things happened in.

Note

Commands you ran in the same worker are deliberately absent. Those belong to the worker, not to the task.

Spunto once let a project point a status command at its forge: a script run in the worker every few seconds whose output could move a task — including all the way to done.

It's gone, and the reason is worth stating, because it looks like a feature to lose. It meant your forge could reach into your machines. Merge a pull request from your phone, and a container you were still reviewing got parked, a task you had never looked at closed itself, and the pool handed that machine to the next task — overwriting the working tree you had merged from. The trigger was elsewhere, the consequence was here, and nothing on screen connected the two.

So a task's state comes from what Spunto started and what you clicked, and from nothing else:

The sessionThe task
is runningrunning
ended badlyfailed
ended cleanlyin-review — and it waits

The same reasoning is why the pull request panel is read-only, and why there's no Merge button on it. Reading doesn't have that power; writing would.

Your forge is still where the work lands, and your accept command is still where a merge happens — the difference is that you trigger it, from the task you're looking at. If a pull request was merged elsewhere, pressing Accept is how you tell Spunto the work is over; the merge command is idempotent on an already-merged branch, or you leave it empty and Accept simply closes the task.

There's no "waiting for your input" state either. It can't be detected honestly yet, and a state that lies is worse than one that's missing.

Note

A live task's state is recomputed when someone reads it, and every 30 seconds in the background otherwise. That floor is what lets a task reach review — and therefore park its container under Stop it — with the dashboard closed.

The pull request panel is served by an optional capability — git.pullRequestForBranch. A project whose repositories are raw clone URLs has nobody to ask, so the section is simply absent rather than empty.

The contract is deliberately thin: draft | open | merged | closed, plus a link, a label, a title, checks, a review verdict and a comment count. No forge-specific fields — every one added here is one every forge has to satisfy.

Everything after state is nullable, and that's meaningful. checks: null means this forge doesn't report them, which is not nothing failed — a repository with no CI hasn't succeeded at anything either, and the panel won't draw a green tick for it.

Two details already paid for on the GitHub side: check runs (Actions, any CI that's an App) and legacy commit statuses (Buildkite, a Jenkins webhook) are two halves that have to be read together, and only each person's latest review counts — summing all of them leaves someone who requested changes and then approved stuck in "changes requested" forever.

A branch pushed to a fork isn't found, and that's stated rather than worked around: finding it would mean sweeping every open pull request of the repository on every poll.

The browser polls slowly (30 s, and once only for a finished task), and the API caches the answer for a few seconds — each response costs several forge calls against a budget the whole organization shares. That's the opposite of the diff, which is never cached, and both are right: a diff describes a machine you can still talk to, a pull request describes a remote object that moves by the minute.

# Accept command
gh pr merge "$SPUNTO_TASK_BRANCH" --squash --delete-branch
 
# Drop command
gh pr close "$SPUNTO_TASK_BRANCH" --delete-branch

These run in the task's own worker — restarted first if review mode had parked it — so they can touch the checkout as well as the forge. Accept requires exit code 0; a failing merge leaves the task in review.

Reusing a worker means putting it back first, and that happens when a task picks it up, not when the previous one lets go — so a machine parked mid-review is never cleaned behind your back.

Spunto handles the git side itself. Before the new branch is checked out, the checkout is reset and cleaned, so whatever the previous task left uncommitted is gone and your agent starts on the base you asked for. Ignored files — node_modules, build caches, .venv — are deliberately kept, because they're what makes reuse worth anything.

Everything else is yours to define with a Reset command: reseed a database, reinstall dependencies, restart a service.

The pool is "workers tagged task, owned by this task's creator, that no live task is holding". A worker owned by another member is never taken, and neither is one without the tag.

Nothing in a task's lifecycle calls GitHub. The only place a forge appears is the clone that populates the worker — so the real prerequisite is a git repository exists in the workspace, not an account is connected. A project can build its own, which is the shortest way to try delegated work on a machine where you have no credentials set up:

# postCreate — a repository and its origin, inside the worker
cd /workspace
if [ ! -d .git ]; then
  git init -q -b main .
  echo ".sandbox-origin.git/" > .gitignore
  git config user.email you@example.com && git config user.name "Sandbox"
  git add -A && git commit -qm "initial"
  git init -q --bare -b main .sandbox-origin.git
  git remote add origin /workspace/.sandbox-origin.git
  git push -q -u origin main
fi
# Agent command — reads the prompt on stdin, makes a real commit on the task branch
sh -c 'prompt=$(cat); { echo "# $SPUNTO_TASK_TITLE"; echo; echo "$prompt"; } > "TASK-$SPUNTO_TASK_ID.md";
       git add -A; git commit -qm "$SPUNTO_TASK_TITLE"; git push -q origin HEAD'

The whole cycle then runs unchanged — worker, branch, session, review, diff, Accept — with no account anywhere. Clear the agent command when you want the real one back.

Tip

Keep the git identity local to the repository and the origin inside /workspace, as above. Both then live in the workspace volume, so a worker recreated on the same volume still finds its branches.

curl -H "Authorization: Bearer spk_..." \
  https://spunto.net/api/orgs/{orgId}/projects/{projectId}/tasks/{taskId}
{
  "id": "k3f9x2...",
  "title": "The Ship button overflows its card on mobile",
  "state": "in-review",
  "branch": "task/the-ship-button-overflows-its-card-a1b2c3",
  "baseBranch": "main",
  "model": "claude-opus-5",
  "workerId": "wkr_abc123",
  "commandId": "cmd_def456",
  "outcomeUrl": "https://github.com/acme/app/pull/262",
  "outcomeLabel": "#262",
  "meta": {},
  "startedAt": "2026-08-30T09:14:02.000Z",
  "completedAt": null
}

outcomeUrl and outcomeLabel are whatever the work produced — a pull request, a merge request, a preview deployment. Nothing in the model is specific to one forge; the criterion for a typed column is the product draws it, and the product draws a link and its label.

commandId is the agent session, readable as an ordinary worker command: its output is there while it runs and after it finishes. Because the session lives in the worker rather than in the API, it survives restarts on our side entirely.

POST   .../projects/{projectId}/tasks                       delegate { prompt, title?, baseBranch?, model?, files? }
GET    /api/orgs/{orgId}/tasks                              every task of the org (?state=&q=&projectId=&limit=&offset=)
GET    .../projects/{projectId}/tasks                       one project's, most recently active first
GET    .../tasks/{taskId}                                   detail
GET    .../tasks/{taskId}/commands                          what the platform ran, oldest first
GET    .../tasks/{taskId}/events                            the session, event by event (?since=&source=)
GET    .../tasks/{taskId}/diff                              what the branch changes, read with git in the worker
GET    .../tasks/{taskId}/diff/patch                        one file's patch (?path=&repo=&oldPath=&untracked=)
GET    .../tasks/{taskId}/pull-request                      what the forge says, if it can be asked
GET    .../tasks/{taskId}/attachments/{attachmentId}        an attachment's bytes
POST   .../tasks/{taskId}/messages                          talk to the session (202)
POST   .../tasks/{taskId}/interrupt                         end the current turn, keep the session (202)
POST   .../tasks/{taskId}/validate                          Accept (202)
POST   .../tasks/{taskId}/cancel                            Drop (202)

Everything that does real work runs as a durable background job — task.run, task.settle, task.validate, task.cancel. None of it dies with an API redeploy, which is why the POSTs that act on an existing task answer 202 instead of waiting. Delegating answers straight away with the queued task.