Observability

OpenTelemetry with zero configuration, resource history, logs, port detection and job runs — built into every worker and deployment service.

The Run side of Spunto covers what happens once your code is live: watching it, debugging it, and re-running one-off tasks against it. This works the same way for development workers and production deployment services.

Tail container stdout/stderr in real time from the dashboard, or fetch the last N lines over the API:

GET /api/orgs/:orgId/projects/:projectId/workers/:workerId/logs?tail=200
GET /api/orgs/:orgId/deployments/:deploymentId/services/:serviceId/logs?tail=200

The dashboard's log panel streams over a WebSocket for live tailing; the tail query param is used for one-off fetches.

Every worker and every deployment service starts with the standard OpenTelemetry variables already set:

OTEL_EXPORTER_OTLP_ENDPOINT=http://spunto-telemetry:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf        # http/json for workers
OTEL_SERVICE_NAME=<service name>                  # deployment services
OTEL_RESOURCE_ATTRIBUTES=spunto.deployment.id=…,spunto.service.id=…

So any OpenTelemetry SDK or auto-instrumentation — Node, Python, Go, Java… — exports its traces, logs and metrics to Spunto without a line of setup. spunto-telemetry is the node's agent, reachable on your service's private network; it forwards to Spunto, which attributes everything to your organization. Both OTLP/HTTP encodings (protobuf and JSON) are accepted.

Your own settings win: set OTEL_SERVICE_NAME or point OTEL_EXPORTER_OTLP_ENDPOINT elsewhere and Spunto steps aside (your OTEL_RESOURCE_ATTRIBUTES are kept, Spunto's ids are appended to them).

Tip

A service created before this existed gets the variables on its next Deploy. A running container keeps the environment it was created with.

Open a deployment and switch to Observability. Filter by service and time range (1 h → 7 days):

  • Headline numbers — traces, error rate, p50 / p95 latency.
  • Resources — CPU, memory and network of each service's container, sampled every 30 s by the node, kept as history across redeploys.
  • Traces — newest first; open one to see its waterfall across services (a Node frontend calling a Python backend is one trace) and the log lines written during it.
  • Logs — what your services emit over OpenTelemetry, filterable by severity and text. (Their raw stdout stays in each service's Logs tab, live.)
  • App metrics — your own counters, gauges and histograms, one sparkline each. Label sets are summed; cumulative counters are shown as reported.

The same data over the API:

GET /api/orgs/:orgId/deployments/:deploymentId/observability/traces?serviceId=&lookbackMinutes=
GET /api/orgs/:orgId/deployments/:deploymentId/observability/traces/:traceId
GET /api/orgs/:orgId/deployments/:deploymentId/observability/logs?search=&minSeverity=&traceId=
GET /api/orgs/:orgId/deployments/:deploymentId/services/:serviceId/metrics?range=1h
GET /api/orgs/:orgId/metrics/names?subjectKind=workload&subjectId=:serviceId
GET /api/orgs/:orgId/metrics?subjectKind=workload&subjectId=:serviceId&metric=shop.orders&range=24h

The repository ships a small shop in examples/otel-demo: a Node storefront (OpenTelemetry SDK) that calls a Python inventory (auto-instrumented), and a loadgen that keeps them busy — three public images, no build. Deploy it into your organization with an API key:

SPUNTO_API=https://your-spunto SPUNTO_TOKEN=spk_… node examples/otel-demo/deploy.mjs <orgId>

About a minute later (the services install their dependencies on first start), its Observability tab fills up.

Every worker and deployment service exposes a live resource snapshot:

GET /api/orgs/:orgId/projects/:projectId/workers/:workerId/stats

Returns CPU% and memory usage from the node's latest sample of the container (every 30 s). Returns 503 if the container isn't running — there's nothing to sample. The history is on the worker page (History) and a deployment's Observability tab.

Any port your code listens on inside a worker is detected automatically (ss -tlnpH is polled inside the container every 10 seconds) and exposed without any forwardPorts configuration:

GET /api/orgs/:orgId/projects/:projectId/workers/:workerId/ports
https://worker-{index}-{id}-{port}.{BASE_DOMAIN}

The Ports panel on the worker page turns every detected port into a clickable link.

Deployment services can define reusable jobs — one-off migrations, cron-style scripts, data exports — and trigger them on demand:

POST /api/orgs/:orgId/deployments/:deploymentId/services/:serviceId/jobs/:jobId/run
GET  /api/orgs/:orgId/deployments/:deploymentId/services/:serviceId/jobs/:jobId/runs

Each run stores its full log output, exit status, start/completion time, and is kept permanently in the run history — searchable and paginated from the dashboard.

Tip

Jobs run in their own temporary container, on the same deployment network as the service, so they can reach other services (e.g. a database) by name.

Alerting, rates computed from cumulative counters, and per-label-set charts are on the roadmap. This page only documents what you can use right now.