HTTP API

September 16, 2026 ยท View on GitHub

The application API is rooted at /api/v2. A running Server's /openapi.json is the authoritative machine-readable schema for request and response bodies. This page records the behavioral contract that an independent Client or Worker must preserve.

curl http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/openapi.json > openapi.json

FastAPI's interactive Swagger and ReDoc pages are intentionally disabled.

Common contract

  • /health and /openapi.json are unauthenticated. Every /api/v2 endpoint uses the one Server-wide Bearer token when authentication is configured.
  • Request schemas reject unknown fields. Response schemas may gain optional fields within v2; existing fields, types, defaults, meanings, and Task states do not change without a new API prefix.
  • Request bodies and complete stored Task data, including progress, are each limited to 1 MiB. Large artifacts do not belong in Task JSON.
  • Public timestamps are UTC RFC 3339 strings. Task states are exactly pending, running, succeeded, failed, and cancelled.
  • Successful 204 responses have no body. Clients must not invent receipt objects for them.

When enabled, authenticate application requests with:

Authorization: Bearer <token>

Application responses under /api/ include Labtasker-Server-Version with the Server package version, including empty successes and handled errors. With authentication enabled, only requests carrying the valid token receive it; otherwise it is public. /health and /openapi.json do not include it.

Clients can observe this header without a preflight request. Older Servers or proxies may omit it; absence is not proof of incompatibility. The Python Client warns on stderr when a known Server version is older, while preserving the operation's result or error. Package version differences do not themselves establish protocol incompatibility.

Discovery

Method and pathSuccess
GET /health200 with {"status":"ok","api_version":"2","database":"ok"} after a real database check. Database failure returns 503 with status and database set to error.
GET /openapi.json200 with the generated OpenAPI schema.

Queue endpoints

Method and pathInputSuccess and constraints
PUT /api/v2/queues/{queue}No body201 when created, 200 when already present; returns {"name": ...}.
GET /api/v2/queuesNone200 with the complete unpaginated Queue array.
DELETE /api/v2/queues/{queue}?cascade=falseBoolean query parameter204. A non-empty Queue requires cascade=true; any running Task still blocks deletion. Successful cascade atomically deletes the Queue and all its Tasks.

Queue names are explicit path components and are not authentication identities.

Task resource endpoints

Method and pathInputSuccess and constraints
PUT /api/v2/queues/{queue}/tasks/{task_id}Task creation object201 on creation; 200 for an identical replay at the same ID; 409 task_id_conflict for a different definition.
GET /api/v2/queues/{queue}/tasks/{task_id}None200 with one Task; 404 task_not_found when absent.
GET /api/v2/queues/{queue}/tasksSelection and pagination query parameters200 with {"items":[...],"next_cursor":...}. Returns one page only.
GET /api/v2/queues/{queue}/tasks/countSelection query parameters200 with {"count":N}, or a grouped page when group_by is supplied.
PATCH /api/v2/queues/{queue}/tasks/{task_id}Non-empty Task update object200 with the Task. Running Tasks reject updates. Object/list fields are complete replacements.
PATCH /api/v2/queues/{queue}/tasks{"filter":...,"changes":...}200 with {"matched":N,"updated":M}. The filter is required; the update is atomic across matching non-running Tasks.
POST /api/v2/queues/{queue}/tasks/{task_id}/progress{"run_id":...,"progress":{...}}204; the matching active run replaces its latest progress snapshot without renewing its lease or changing Task status.
POST /api/v2/queues/{queue}/tasks/{task_id}/cancelNo body200 with the cancelled Task. Accepts pending/running and is idempotent for cancelled.
POST /api/v2/queues/{queue}/tasks/{task_id}/requeueNo body200 with the pending Task. Accepts pending/failed/cancelled, resets attempt and last error.
DELETE /api/v2/queues/{queue}/tasks/{task_id}None204; idempotent when absent. Running Tasks reject deletion.

Creation body

Every field is optional because the Server expands the same defaults as the Python API:

{
  "name": null,
  "args": {},
  "metadata": {},
  "priority": 0,
  "max_attempts": 3,
  "routes": ["default"]
}

The Task ID is the resource identity and must match t_[A-Za-z0-9_-]{12}. Replaying the same normalized definition returns the Task's current representation. JSON object-key order, route input order, and an omitted default versus the same explicit default do not change that identity.

List and count queries

Task listing accepts:

status, name, name_fuzzy, filter, order_by, descending, limit, cursor

Count accepts status, name, name_fuzzy, filter, and optional group_by with grouped limit/cursor. Selectors are combined with AND. limit is 1 to 1000 and defaults to 100. A non-null next_cursor must be reused with the same Queue, selectors, filter, order field, and direction. The cursor is opaque. See Query language for filter syntax.

name_fuzzy searches names using case-insensitive subsequences. Each whitespace-separated word must match; words may occur in any order. Empty or whitespace-only search adds no restriction. Non-empty search excludes unnamed Tasks. Matching happens on the Server before pagination, uses literal punctuation, and preserves ordering. Exact name and filter equality remain available.

Grouped counts

GET /api/v2/queues/{queue}/tasks/count?status=pending&group_by=routes,status returns, for example:

{
  "group_by": ["routes", "status"],
  "count": 2,
  "items": [
    {"key": {"routes": "a", "status": "pending"}, "count": 2},
    {"key": {"routes": "b", "status": "pending"}, "count": 1}
  ],
  "next_cursor": null
}

One selected Task in this example supports both a and b. Selection applies before route expansion. count is the deduplicated complete Task total; groups can overlap. Task grouping supports only routes, status, or their combination in either order, with one comma-separated group_by parameter and no spaces. No arbitrary aggregation expressions or separate route records are introduced.

Groups sort lexically by values in the requested dimension order. Only nonzero groups appear. limit defaults to 100 groups, with range 1โ€“1000. Cursors bind the resource, Queue, complete selection and ordered dimensions; page size may change. Malformed or mismatched cursors return 422 invalid_cursor. Count and items share a read transaction, but subsequent pages read current data.

Empty grouping, whitespace, unsupported or repeated dimensions, repeated group_by parameters, and pagination without grouping return 422 invalid_request. Ungrouped responses remain {"count":N}. A Client requesting grouping must reject an old Server's scalar response rather than treating it as grouped data.

Update body

A Task update contains at least one of:

{
  "name": "new name",
  "args": {},
  "metadata": {},
  "priority": 10,
  "max_attempts": 5,
  "routes": ["robotwin"],
  "result": {}
}

Every supplied object/list is a complete replacement. Identity, status, attempt, last error, run ownership, and timestamps are Server-owned. Bulk update counts only rows that match and remain non-running; a concurrent claim either sees all new values or excludes that Task. Validation failure for one matched non-running Task rolls back the complete batch.

Worker observations

All paths below require an existing Queue and the ordinary application token.

Method and pathInputSuccess
GET /api/v2/queues/{queue}/workersfilter, limit, cursor200 with {"items":[...],"next_cursor":...}, ID ascending, default 100/max 1000.
GET /api/v2/queues/{queue}/workers/countfilter, optional group_by, grouped limit/cursor200 with scalar or grouped counts. Dimensions: route, status, or both in either order.
PUT /api/v2/queues/{queue}/workers/{id}Complete {"route":"sdxl","status":"busy","task_id":"t_ABCDEFGHIJKL","metadata":{"hostname":"node-7"}}204; creates or renews the observation.
POST /api/v2/queues/{queue}/workers/{id}/telemetry{"telemetry":{"gpu_utilization":0.75}}204; replaces telemetry without renewing the observation.
DELETE /api/v2/queues/{queue}/workers/{id}None204, including an absent Worker in an existing Queue.

Each observation exposes id, queue, route, status, nullable task_id, metadata, nullable telemetry and telemetry_updated_at, last_seen_at, and expires_at. IDs match w_[A-Za-z0-9_-]{12} and identify one loop invocation across successive Tasks. Fixed fields plus nested metadata.* and telemetry.* paths support the existing filter grammar; Task-only paths are rejected. Worker count pagination and filtering follow the grouped-count rules above. Grouping remains limited to route and status.

A report requires the three original fields; metadata is a strict JSON object defaulting to {} for older reporters. task_id accepts null or a valid Task ID without a Task lookup. Updating an existing instance to another route returns 409 worker_route_conflict. idle means waiting for work; busy begins at confirmed claim and covers execution, reporting and cleanup, even after Task completion. The bundled Worker reports null when idle and its Task ID when busy.

Telemetry is a strict, user-defined JSON object with no built-in GPU or platform schema. Reporting is synchronous, one request per call, and complete replacement. The Server timestamps accepted telemetry but does not update last_seen_at, renew expiry, create an absent/expired Worker, retain history, or affect execution. An absent or expired Worker returns 404 worker_not_found.

The Server assigns both timestamps and expires each accepted observation after 300 seconds. Bundled Workers report every 60 seconds and on activity changes through an independent reporter. Only unexpired rows appear in reads; expired rows are deleted at startup and on the lease-scan cadence. These are latest observations, with no history, process-control or Task-ownership effect. An advisory Task reference can be terminal, deleted or stale. Counts are approximate.

Observation failures never stop or block the Task loop or affect its failure guard. Once the loop has independently decided to exit, it waits at most one second for reporter shutdown and best-effort withdrawal. Crashes and failed withdrawals fall back to expiry. A delayed PUT can recreate a withdrawn row; there are no tombstones. Worker observations do not block Queue deletion and are removed with the Queue. Reports never recreate a missing Queue.

Worker protocol

These endpoints let an independent executor implement the same lease/fencing contract as the bundled Workers. They are not public convenience methods on the Python Client.

Method and pathBodySuccess
POST /api/v2/queues/{queue}/tasks/claim{"route":...,"run_id":...}200 with Task, run_id, and lease_expires_at; 204 when no compatible pending Task exists.
POST /api/v2/queues/{queue}/tasks/{task_id}/heartbeat{"run_id":...}200 with the renewed lease_expires_at.
POST /api/v2/queues/{queue}/tasks/{task_id}/progress{"run_id":...,"progress":{...}}204; replaces the latest snapshot and does not renew the lease.
POST /api/v2/queues/{queue}/tasks/{task_id}/complete{"run_id":...,"result":{...}}204; completes the matching active run as succeeded.
POST /api/v2/queues/{queue}/tasks/{task_id}/fail{"run_id":...,"error":{"type":...,"message":...,"traceback":...}}204; charges the failure and retries or fails according to the Task budget.
POST /api/v2/queues/{queue}/tasks/{task_id}/unclaim{"run_id":...}204; returns the matching run to pending without an error payload.

run_id matches r_[A-Za-z0-9_-]{12} and is a private per-claim ownership token, not a Run resource. Heartbeat and terminal actions require the currently active ID. A stale process cannot mutate a newer claim.

Claim matches the supplied route exactly and case-sensitively. Among compatible pending Tasks, higher priority is claimed first, followed by stable pending order. A healthy run renews its five-minute lease with heartbeat; the lease is not an execution timeout.

Terminal actions are effect-idempotent for the most recently finalized run and same action. A contradictory action or stale run conflicts. The Server exposes no /runs collection or execution-history API.

Task representation

Ordinary get/list/action responses contain exactly these required v2 fields:

id, queue, status, name, args, metadata, priority, attempt, max_attempts,
routes, result, progress, progress_updated_at, progress_attempt, last_error,
last_route, created_at, updated_at, started_at, finished_at

progress is null until the current attempt reports a strict JSON object. progress_updated_at and progress_attempt are Server-owned and are null at the same time. Reports completely replace the object; they do not merge it into the final result. Run completion, failure, expiry and cancellation retain the last snapshot, while the next successful claim clears it. Dynamic progress.* paths are available to Task filters.

Active run_id and lease expiry appear only in Worker protocol responses, never in the ordinary Task resource.

Errors

Every documented application error uses one envelope:

{
  "error": {
    "code": "stale_run",
    "message": "This run is no longer active.",
    "details": {}
  }
}

Branch on error.code, not message. The stable status classes are:

HTTP statusMeaning and representative codes
401Missing or incorrect token: unauthorized.
404Missing Queue or Task: queue_not_found, task_not_found.
409State or ownership conflict: task_id_conflict, task_running, stale_run, run_finalized, lifecycle/update conflicts.
413Body exceeds 1 MiB: request_too_large.
422Invalid request, Task, filter, update, identifier, or complete stored data: for example invalid_task, invalid_filter, invalid_update, task_data_too_large.
503Retryable Server-side operational failure such as database_busy; still uses the same envelope.

Malformed FastAPI/Pydantic native errors are converted to this envelope; Server tracebacks and database details are never part of the public response.

Client retry boundary

Safe reads and representation-idempotent Task creation can be retried with the same inputs. Lifecycle, update, and delete calls should not be retried blindly after an uncertain response because another actor may have changed the resource. Inspect current state before deciding.

Workers use a separate reliability policy: they keep heartbeat active while retrying one terminal report, and run_id fencing prevents a stale retry from overwriting a newer execution.