Skip to main content

The ChronosOS Server (chronos serve)

chronos serve starts ChronosOS — the Chronos control plane. It is a single, self-contained HTTP server that exposes everything you need to operate agents in production: a REST API over sessions, checkpoints, traces, schedules and human-in-the-loop approvals; a live server-sent-events (SSE) firehose; Prometheus metrics; health/readiness probes; and an interactive Swagger UI.

The control plane is stateless — all durable state lives in the configured storage backend. You can therefore run many replicas behind a load balancer, and any replica can serve any request.

:::note At a glance

  • Default listen address: :8420
  • REST API mounted under /api/*
  • Swagger UI at /swagger/, OpenAPI JSON at /swagger/doc.json
  • Auth is opt-in (default: none) — see Authentication & Authorization
  • Hardened by default: timeouts, body limits, panic recovery, CORS, rate limiting, graceful shutdown :::

Starting the server​

From the CLI​

# Start on the default address (:8420)
chronos serve

# Bind a custom address
chronos serve :9000
chronos serve 127.0.0.1:8420

The server reads its storage configuration from environment variables (see Storage and Configuration). On start it logs the resolved listen address, the storage backend, and whether auth is enabled.

ChronosOS itself has no concept of "the agent roster" — agents are a Go/YAML concept from the sdk/agent package, not something the control plane runs by itself. chronos serve does, however, optionally load an agents YAML file (the same -c <file> / CHRONOS_CONFIG resolution as chronos run/repl, falling back to .chronos/agents.yaml if present) so that any agent marked durable: true gets its compiled graph registered with the dashboard — see YAML Agents & the Dashboard. A config file is entirely optional: with none found, serve starts exactly as before, with zero agents registered.

Stop it with Ctrl-C (SIGINT) or SIGTERM — the server drains in-flight requests before exiting (see Graceful shutdown).

From the SDK​

For embedding the control plane in your own binary, use the os package (package name chronosos, to avoid colliding with the standard library os package). The functional-options constructor lets you enable auth, tune the hardening defaults, and attach a scheduler, approval service, or a store-backed rate limiter:

package main

import (
"context"
"database/sql"
"log"

chronosos "github.com/spawn08/chronos/os"
"github.com/spawn08/chronos/os/auth"
"github.com/spawn08/chronos/os/middleware"
"github.com/spawn08/chronos/storage/adapters/postgres"
)

func main() {
ctx := context.Background()

store, err := postgres.New("postgres://user:pass@db:5432/chronos")
if err != nil {
log.Fatal(err)
}

// Store-backed limiter so limits are shared across replicas — shares the
// same *sql.DB as your storage backend.
db, err := sql.Open("postgres", "postgres://user:pass@db:5432/chronos")
if err != nil {
log.Fatal(err)
}
limiter := middleware.NewSQLLimiter(db, middleware.DialectPostgres)
if err := limiter.Migrate(ctx); err != nil {
log.Fatal(err)
}

srv := chronosos.NewWithOptions(":8420", store,
// Opt in to JWT auth (see the Authentication guide)
chronosos.WithJWTAuth(auth.JWTConfig{
Secret: "…", // HS256 shared secret
Issuer: "https://issuer.example.com",
Audience: "chronos",
}),
chronosos.WithRateLimiter(limiter),
)

if err := srv.Start(ctx); err != nil {
log.Fatal(err)
}
}

The simplest form, chronosos.New(addr, store), applies all hardening defaults with no authentication — suitable for a trusted network or local development.

Configuration & options​

Every hardening behaviour has a sensible default and a corresponding SDK option.

OptionDefaultPurpose
WithJWTAuth(auth.JWTConfig{…})offEnable JWT bearer-token auth (HS256 / RS256 / JWKS). See Authentication.
WithAPIKeyAuth(auth.APIKeyConfig{…})offEnable X-Api-Key header auth with optional per-key quotas.
WithRBAC(true)offEnforce roles on /api/* — reads need viewer, mutations need user. Effective only with auth enabled (CHRONOS_RBAC).
WithSwagger(false)Swagger onDisable the Swagger UI and OpenAPI spec (CHRONOS_SWAGGER).
WithCORS(cfg) / WithoutCORS()CORS onConfigure or disable the CORS middleware.
WithRateLimit(cfg) / WithoutRateLimit()rate limit onConfigure or disable request rate limiting.
WithRateLimiter(middleware.NewSQLLimiter(db, dialect))in-memoryUse a store-backed limiter (backed by a shared *sql.DB) so limits are shared across replicas.
WithTimeouts(read, readHeader, write, idle)see belowOverride the server timeouts.
WithMaxBodyBytes(n)1 MiBMaximum request body size (also caps header size).
WithScheduler(sched)in-process default/api/schedules works out of the box against an in-memory scheduler; use this to swap in a store-backed one (see Running multiple replicas).
WithApproval(svc)in-process default/api/approval/* works out of the box against an in-memory service; use this to swap in a store-backed one.
WithGraphs(dashboard.GraphRegistry{…})none registeredRegister compiled graphs (by agent id) so the dashboard and /api/dashboard/runs can start/resume/time-travel their sessions. chronos serve populates this automatically for durable: true YAML agents — see YAML Agents & the Dashboard.

Hardening defaults​

Applied automatically by both chronos serve and chronosos.New:

SettingDefault
Read timeout30s
Read-header timeout10s
Write timeout30s (cleared for SSE streams)
Idle timeout120s
Max header + body size1 MiB
Panic recoveryon (returns 500, logs stack)
CORSon
Rate limitingon
Graceful shutdownon (SIGTERM / SIGINT)

:::tip Long-lived streams The /api/events/stream SSE endpoint is exempt from the write timeout — the handler clears the write deadline for that connection so the stream stays open indefinitely. See Streaming & SSE. :::

Middleware stack​

Every request passes through the middleware chain in this fixed order. The outermost layer runs first on the way in and last on the way out:

request
→ recovery (catch panics → 500, never crash the server)
→ logging (structured request/response log line)
→ CORS (preflight + response headers)
→ rate limit (429 when the bucket is exhausted)
→ auth (401/403 — bypassed for health, metrics, swagger)
→ route handler

Ordering rationale:

  • Recovery first so it wraps everything, including the logger.
  • Logging before CORS/auth so rejected requests are still recorded.
  • Rate limit before auth so unauthenticated floods are shed cheaply.
  • Auth last so only well-formed, non-throttled requests reach principal resolution — and so tenant scoping is available to the handler.

:::note Always-public paths /healthz, /health, /health/live, /health/ready, /metrics, and everything under /swagger* bypass the auth middleware regardless of the configured auth mode. This keeps liveness/readiness probes and metrics scraping working without credentials. Because Swagger is reachable anonymously, disable it on hardened production servers with CHRONOS_SWAGGER=false (see Authentication). :::

:::note Role enforcement By default the auth middleware only checks that a credential is valid, not what role it carries. Set CHRONOS_RBAC=true (or WithRBAC(true)) to also gate /api/* routes by role — reads require viewer, mutations require user. See Roles & RBAC. :::

Getting a credential quickly​

Once you've turned on auth (CHRONOS_AUTH=apikey or CHRONOS_AUTH=jwt), mint a working credential without writing or signing anything by hand:

export CHRONOS_AUTH=apikey
chronos auth token --role admin
# prints a CHRONOS_API_KEYS entry and a ready-to-use curl example

export CHRONOS_AUTH=jwt
export CHRONOS_JWT_SECRET=some-shared-secret
chronos auth token --role admin --ttl 12h
# prints a signed HS256 JWT and a ready-to-use curl example

chronos auth token reads the same CHRONOS_AUTH/CHRONOS_JWT_SECRET/ CHRONOS_API_KEYS environment the server itself reads, so the credential it mints is guaranteed to match whatever mode chronos serve will start with.

Health & readiness​

Four health endpoints support container orchestrators and load balancers. All return JSON {"status":"ok"} (HTTP 200) when healthy and are always unauthenticated.

EndpointSemanticsTypical use
GET /healthzProcess is upLegacy / general health check
GET /healthProcess is upAlias of /healthz
GET /health/liveLiveness — the process is runningKubernetes livenessProbe
GET /health/readyReadiness — storage reachable, ready for trafficKubernetes readinessProbe

Liveness answers "should this pod be restarted?"; readiness answers "should this pod receive traffic?". A pod that is alive but not ready (e.g. its database is briefly unreachable) is pulled from the load-balancer rotation without being killed.

# Kubernetes probe wiring
livenessProbe:
httpGet: { path: /health/live, port: 8420 }
readinessProbe:
httpGet: { path: /health/ready, port: 8420 }

Metrics​

GET /metrics exposes Prometheus-format metrics (request counts, latencies, in-flight requests, and Go runtime stats). It is unauthenticated so a Prometheus scraper needs no credentials:

# prometheus.yml
scrape_configs:
- job_name: chronos
static_configs:
- targets: ["chronos:8420"]

The metrics.Registry​

/metrics is served by a *metrics.Registry from the os/metrics package. chronosos.NewWithOptions creates one automatically (metrics.NewRegistry()) and exposes it as srv.Metrics; it pre-registers the Chronos counters, gauges, and histogram used throughout the control plane (chronos_agent_runs_total, chronos_tool_calls_total, chronos_tokens_used_total, chronos_model_calls_total, chronos_errors_total, chronos_active_sessions, chronos_model_latency_seconds, chronos_tool_latency_seconds).

To feed it from agent execution (agents run outside the control plane), add hooks.NewPrometheusHook(srv.Metrics) to the hook chain of the agents you run — see Middleware & Hooks. You can also register your own series directly:

requests := srv.Metrics.Counter("myapp_requests_total", "Total app requests")
requests.Inc(map[string]string{"route": "/checkout"})

Registry.Handler() returns the http.Handler mounted at /metrics; you only need it yourself if you are serving metrics from a mux that isn't ChronosOS.

Exporting to an OTLP collector​

For push-based pipelines (e.g. an OpenTelemetry Collector that forwards to a remote-write endpoint), export the registry over OTLP/HTTP with metrics.NewOTLPExporter, instead of — or in addition to — letting Prometheus scrape /metrics:

package main

import (
"context"
"log"
"time"

chronosos "github.com/spawn08/chronos/os"
"github.com/spawn08/chronos/os/metrics"
)

func main() {
ctx := context.Background()

srv := chronosos.New(":8420", store) // store: your storage.Storage

// POSTs the registry snapshot to "<endpoint>/v1/metrics" as OTLP/JSON.
exporter := metrics.NewOTLPExporter("http://otel-collector:4318")

go func() {
ticker := time.NewTicker(15 * time.Second)
defer ticker.Stop()
for range ticker.C {
if err := exporter.Export(ctx, srv.Metrics); err != nil {
log.Printf("otlp export: %v", err)
}
}
}()

if err := srv.Start(ctx); err != nil {
log.Fatal(err)
}
}

NewOTLPExporter(endpoint) builds an exporter that targets "<endpoint>/v1/metrics"; passing an empty endpoint makes Export a no-op (exporter.Enabled() reports false), which keeps offline runs and tests safe. Chain WithHeader(key, value) to attach auth headers (e.g. an Authorization bearer token) and WithHTTPClient(c) to override the default 10s-timeout client. Export marshals the registry's counters, gauges, and histograms into a cumulative-temporality ExportMetricsServiceRequest and POSTs it once per call — call it on your own schedule (as above) since the exporter does not run a background loop.

Streaming & SSE​

GET /api/events/stream is a long-lived text/event-stream connection that pushes graph and run events as they happen:

  • ?session=<id> scopes the stream to a single session.
  • No query parameter subscribes to the firehose — every session's events on this replica.
# Follow one session
curl -N http://localhost:8420/api/events/stream?session=sess-123

# Firehose (all sessions)
curl -N http://localhost:8420/api/events/stream

Browser clients use the standard EventSource API:

const es = new EventSource("/api/events/stream?session=sess-123");
es.onmessage = (e) => console.log(JSON.parse(e.data));

The write deadline is cleared for this handler so the connection never times out mid-stream. For the underlying broker and event types, see Streaming & SSE.

Multi-tenancy​

When auth is enabled, every authenticated principal carries a TenantID claim. The server derives the tenant from the principal and scopes every storage operation to it — so a caller can only ever see sessions, traces, checkpoints, schedules, and approvals belonging to their own tenant. This makes the API IDOR-safe: passing another tenant's session_id returns no data rather than leaking it.

With auth disabled, all requests run under the single DefaultTenant. See Multi-Tenancy for the storage-layer model.

Graceful shutdown​

On SIGTERM or SIGINT the server:

  1. Stops accepting new connections.
  2. Lets in-flight requests finish (bounded by a shutdown timeout).
  3. Closes the scheduler, approval service, and storage handles.
  4. Exits 0.

This makes rolling deploys safe — Kubernetes sends SIGTERM, and the pod drains before its grace period elapses.

Running multiple replicas​

The control plane is stateless, so horizontal scaling is a matter of pointing every replica at the same shared storage and using store-backed components where per-request coordination is needed:

ConcernSingle replicaMultiple replicas
Sessions / checkpoints / tracesStorage backendSame shared backend
Rate limitingIn-memory (default)WithRateLimiter(middleware.NewSQLLimiter(db, dialect)) — shared buckets
SchedulerWithSchedulerStore-backed scheduler with leasing so a cron job fires once across the fleet
ApprovalsWithApprovalStore-backed approval service — any replica can resolve a pending approval
SSE firehoseAll eventsEach replica streams the events it processes; subscribe per session, or fan-in downstream

:::warning SQLite is single-node Use PostgreSQL (or another networked backend) for multi-replica deployments. SQLite is file-local and cannot be shared safely across pods. :::

Production deployment​

A complete, opinionated production example — Postgres storage, JWT/JWKS auth, TLS termination at the ingress, HPA, and probes — lives in deploy/production/ in the repository. For Kubernetes and Helm specifics see Kubernetes & Helm; for the full request-per-endpoint contract see the REST API Reference.

See also​