Observability
NimbleBrain emits vendor-neutral telemetry: OpenTelemetry traces over OTLP, structured JSON logs, and a Prometheus metrics endpoint. None of it requires infrastructure to run — with no collector configured, traces are simply not exported and trace ids still flow through logs for correlation. This page covers what to wire up when you run NimbleBrain in production.
Tracing (OpenTelemetry)
Section titled “Tracing (OpenTelemetry)”The runtime wraps its hot paths in OpenTelemetry spans and exports them over OTLP/HTTP. It depends only on the open OTel and W3C tracecontext primitives, so you point it at any OTLP-compatible collector (the OpenTelemetry Collector, Grafana Tempo, Jaeger, Honeycomb, and so on).
The runtime produces these spans today:
| Span | Covers |
|---|---|
agent.turn |
One engine run (a full agent turn) |
llm.call |
A single model stream |
tool.dispatch |
An MCP tool dispatch |
| HTTP server span | The inbound request; continues an upstream traceparent if present |
Spans nest automatically and carry the verified request identity (user_id, workspace_id, and either conversation_id for a chat or run_id for a scheduled run) plus the boot-time tenant_id. They never carry the user’s display name, email, secrets, prompts, tool arguments or results, or file contents.
Enabling export
Section titled “Enabling export”Set OTEL_EXPORTER_OTLP_ENDPOINT to your collector’s base URL. When it is unset, nothing is exported (the default — local bun run dev and OSS checkouts need no infrastructure).
environment: - OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318 - NB_SERVICE_NAME=nimblebrain-runtime - NB_TENANT_ID=acme| Variable | Default | Purpose |
|---|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT |
unset | Collector base URL. Unset = no export; trace ids still exist for log correlation |
NB_SERVICE_NAME |
nimblebrain-runtime |
Service name stamped on every span and log line |
NB_TENANT_ID |
unset | This deployment’s tenant. A boot-time constant stamped on every span and log line — never read from a request header |
Propagation
Section titled “Propagation”Outbound calls that should extend a trace — minting a service token, fetching an authenticated remote MCP server — inject the active W3C traceparent so the trace continues across the hop. Inbound, the HTTP middleware continues any upstream traceparent it receives. No configuration is required; this is on whenever a span is active.
Structured logs
Section titled “Structured logs”By default the runtime logs pretty, colored lines to stderr for the local dev loop. Set NB_LOG_FORMAT=json to switch to one structured JSON object per line — this is what you want for deployed pods (the platform chart sets it for you).
environment: - NB_LOG_FORMAT=jsonEach JSON line is auto-enriched with:
| Field | Source |
|---|---|
service |
NB_SERVICE_NAME |
tenant_id |
NB_TENANT_ID (boot-time, omitted in dev) |
correlation_id |
The active trace id — pivot a log straight to its trace |
user_id / workspace_id |
The verified request identity (omitted when not in a request) |
conversation_id / run_id |
What the work belongs to — a chat carries the first, a scheduled run the second, never both |
Identity comes only from the verified request context, never the wire, and never includes the human display name. Logs stay on stderr, so Kubernetes captures them alongside stdout and stdout stays clean for JSON-RPC and piped output.
Error reporting (Sentry, optional)
Section titled “Error reporting (Sentry, optional)”The OpenTelemetry pipeline above is the system of record for traces and logs. Sentry is an optional, additive error sink on top of it — turn it on when you want uncaught exceptions and unhandled promise rejections grouped, deduplicated, and alertable with full stack traces, without scraping logs for them.
It is off by default and switched explicitly — set NB_SENTRY_ENABLED=true (with NB_SENTRY_DSN) to turn it on. Enablement is never inferred from DSN presence, so a stray SENTRY_DSN in the environment can’t silently wire a process into Sentry. When off, the SDK is never even loaded (the OSS, local-dev, and test default). In multi-tenant deployments the chart sets NB_SENTRY_ENABLED per tenant from runtime.config.sentry.enabled, so one tenant can run with Sentry on and another off. These NB_SENTRY_* vars mirror the web client — our own knobs, not Sentry’s auto-read SENTRY_*.
environment: - NB_SENTRY_ENABLED=true - NB_SENTRY_DSN=https://<key>@<org>.ingest.sentry.io/<project> - NB_SENTRY_ENV=production| Variable | Default | Purpose |
|---|---|---|
NB_SENTRY_ENABLED |
false |
Explicit on/off switch. Must be true to enable — never inferred from the DSN |
NB_SENTRY_DSN |
unset | The project DSN (required when enabled) |
NB_SENTRY_ENV |
unset | Environment tag (e.g. staging, production) — one project, discriminated by this tag |
NB_SENTRY_TRACES_SAMPLE_RATE |
0 |
Sentry performance tracing. 0 = errors only (recommended) |
Release is the runtime’s existing NB_VERSION → NB_BUILD_SHA build identity (the same values /v1/bootstrap reports), so events group by deploy with no extra var.
Prometheus metrics
Section titled “Prometheus metrics”The runtime exposes a Prometheus exposition endpoint at the bare path GET /metrics:
curl http://platform:27247/metricsA few properties worth knowing:
- It is served at the bare
/metricspath, not under/v1. The bundlednimblebrain-webCaddy proxy only forwards/v1/*, so/metricsis not reachable through the public web container — it is scraped in-cluster. - It is unauthenticated. It carries no per-user data and is meant to be reachable only inside the cluster.
In Kubernetes, point a ServiceMonitor (or your scrape config) at the platform service’s /metrics path. Do not expose it to the internet.
Task run metrics
Section titled “Task run metrics”| Metric | Type | Labels | What it says |
|---|---|---|---|
nb_task_runs_total |
counter | status |
Task runs recorded, by status: success, degraded, failure, timeout, cancelled, skipped. See run history for what each means. |
degraded counts runs that finished with a tool call that failed and was never retried to success: an email that was not sent, or a record that was not written. Alert on the ratio of failure to all runs. Leave degraded out of a paging alert and read it in the run list: most degraded runs are a task acting on something its user deleted, such as a draft or message that no longer exists, which the user fixes and on-call cannot. There is no task or workspace label: both are unbounded, and the run record names the task.
Notification poll metrics
Section titled “Notification poll metrics”An app can declare an outbox — a resource the runtime reads on a schedule for facts the app recorded that nobody asked for. That poll is a standing background cost, so it is metered by what it costs rather than by what it carried:
| Metric | Type | Labels | What it says |
|---|---|---|---|
nb_notifications_pulled_total |
counter | source |
Envelopes read off an outbox and admitted to the inbox. |
nb_notifications_poll_seconds |
histogram | source |
Duration of one outbox read, including any reconnect it had to do first. |
nb_notifications_poll_reconnects_total |
counter | source |
Reads that found the connection idle-closed and re-established it. The poll budget is sized on the assumption that this costs three requests; a rate near zero means the assumption is pessimistic. |
nb_notifications_poll_deferred_total |
counter | — | Polls pushed to a later tick because a workspace’s notifications.poll.budgetPerMinute was spent. Sustained nonzero means the budget is smaller than the number of outboxes installed. |
nb_notifications_truncated_total |
counter | source |
Reads answered truncated — the app dropped undelivered events to stay under its own cap, so the inbox has a gap. Any nonzero rate is real data loss and belongs on a dashboard. |
source is the connector’s MCP source name, bucketed to other for a name outside the label charset. There is no workspace label: one runtime serves one tenant, so the scrape already attributes it, and a workspace label would be unbounded.
Credential store metrics
Section titled “Credential store metrics”The states of sealed secrets that should page someone rather than sit in the workspace log:
| Metric | Type | Labels | What it says |
|---|---|---|---|
nb_credential_seal_failures_total |
counter | reason |
Stored secrets that failed to become a usable value, one per audit.credential_seal_failure line. reason is malformed, unknown_kid, auth_failed, no_sealer, plaintext_refused or reseal_skipped. |
nb_credential_store_plaintext_accepted |
gauge | — | 1 when a sealing key is configured and the boot sweep could not prove every secret sealed, so plaintext files are still accepted. 0 after a clean sweep, and 0 with no sealing key, where plaintext is the configuration. Set once per boot. |
nb_credential_store_sealed |
gauge | — | 1 when a sealing key is configured, 0 when secrets are plaintext files. Alert on 0 for a deployment that is meant to be sealed. Set once per boot. |
reason is the only label. A key name would disclose which vendors a deployment uses, and a workspace or user id is unbounded; the audit line beside each increment names the scope, key and wanted key id.
Alert on the counter’s value (nb_credential_seal_failures_total > 0), not on increase() or rate(). The boot sweep runs before the first scrape, so a failure it finds is already in the first sample and never shows as an increase. A restart resets the counter and the next sweep counts the failure again.
Background loops
Section titled “Background loops”Two loops run on timers inside the runtime process: the connection revalidator, which re-checks connectors whose upstream credential can lapse without a transport error (NB_CONNECTION_REVALIDATE_INTERVAL_SECONDS), and the notification poller, which reads declared outboxes (notifications.poll).
Both are per-process and hold their state in memory, so both carry the same prerequisite:
To confirm the poller found anything after a restart, look for [notifications] targets=N — it is logged on the first sweep and again whenever the count changes, so targets=0 on a pod where an app declaring an outbox is installed means the declaration is not reaching the poller, which is a different fault from an outbox that is being read and is empty.
Health and readiness probes
Section titled “Health and readiness probes”The platform serves a liveness/readiness endpoint at GET /v1/health. It answers 200 with a fixed body once the server is serving:
curl http://platform:27247/v1/health{ "status": "ok" }The Docker image already wires this into its HEALTHCHECK (curl -f http://localhost:27247/v1/health), and Compose gates the web service on it via depends_on: condition: service_healthy. In Kubernetes, use the same path for both livenessProbe and readinessProbe.
The endpoint is unauthenticated, and the web container proxies every /v1/* path, so anyone who can reach the web UI can reach it. That is why the body carries nothing else. The web client also calls it before sign-in, to tell a server that is down from a session that has expired.
To see more than liveness:
| To confirm | Read |
|---|---|
| Which build is live | version and buildSha on the authenticated GET /v1/bootstrap, also shown on the About page in organization settings |
| Whether connectors are up | The nb_connector_unhealthy gauge on /metrics: 1 for each connector that is currently down, and no series for a healthy one |