Skip to content

Core Concepts

NimbleBrain’s core value is composing multiple MCP servers into a single, orchestrated workspace. Rather than connecting to one MCP server at a time, NimbleBrain aggregates tools from every installed app into a unified namespace, layers prompts to give the agent awareness of all available capabilities, and filters tool access per-task through skills.

┌──────────────────────────────────────────────────┐
│ ToolRegistry │
│ │
│ ┌─────────┐ ┌─────────┐ ┌──────────────────┐ │
│ │ App 1 │ │ App 2 │ │ App 3 │ │
│ │ (HTTP) │ │ (HTTP) │ │ (SSE) │ │
│ │ 3 tools │ │ 5 tools │ │ 2 tools │ │
│ └─────────┘ └─────────┘ └──────────────────┘ │
│ │
│ Unified namespace: 10 app tools + system tools │
└──────────────────────────────────────────────────┘
│ │
▼ ▼
Skill filtering Agent Engine
(allowed-tools) (agentic loop)

Three mechanisms make this composable:

Mechanism What it does
Tool aggregation The ToolRegistry collects tools from every MCP source (remote HTTP/SSE servers, and the platform’s own in-process ones) and presents them as a flat namespace. The agent doesn’t know or care which server owns a tool.
Skill-scoped filtering Each skill declares allowed-tools glob patterns. When a skill matches, only tools matching those patterns (plus system tools) are visible to the agent. This scopes the agent’s capabilities per task.
4-layer prompt composition System prompts are assembled from: (1) identity, (2) core context, (3) installed app metadata with trust scores, and (4) the matched skill’s prompt. Every layer adds awareness without the agent needing to discover it.

Every chat message triggers an agentic loop in the AgentEngine. The loop repeats until the model produces a response with no tool calls, or a limit is reached.

User message
│
▼
┌─────────────────┐
│ Call Claude │◄─────────────────┐
│ (with tools) │ │
└────────┬────────┘ │
│ │
┌────┴────┐ │
│ Tool │ Yes │
│ calls? │──────► Execute tools │
│ │ in parallel ───┘
└────┬────┘
│ No
▼
Return response

Each pass through the loop is one iteration. Defaults:

Setting Default Hard cap
Max iterations 25 50
Max input tokens 500,000 —
Max output tokens the model’s catalog output limit (16,384 for a model the catalog lacks) —
models.default anthropic:claude-sonnet-4-6 —

The engine stops for one of these reasons:

  • complete — the model responded without requesting any tool calls
  • max_iterations — the iteration limit was reached
  • max_input_tokens — the run has an input-token cap (a task’s Max Input Tokens), and the next model call would pass it. Before each call the engine projects that call’s input as the larger of its estimate of the prompt and the previous call’s reported input, so the run ends within the cap unless the estimate undercounts. The cap counts cache reads, so it bounds tokens, not cost. Whatever the run produced before the stop is kept. maxInputTokens in nimblebrain.json is different: it bounds the context of one model call, not the run.
  • spend_limit — the run names spend accounts (a task’s Token budget is two of them, one per cap), and what is left in one cannot pay for the next model call. Before each call the runtime projects the call’s input as for max_input_tokens, lowers the call’s output ceiling to what every account can pay for, and sets that much aside until the call returns; runs naming the same account share one balance, so they cannot together spend past it. The run stops when an account cannot cover the call’s input, or would leave it only a few hundred output tokens. Whatever the run produced before the stop is kept.
  • cancelled — the run was stopped before it finished: the user pressed Stop, a task run was cancelled or timed out, or the runtime aborted it (for example on shutdown). The conversation keeps what the turn did before the stop.

All tool calls within a single iteration run in parallel via Promise.all().

Before the engine loop starts, the Runtime orchestrates the composition described above:

  1. Resolve or create a conversation
  2. Match the user message to a skill (triggers, then keywords)
  3. Compose the system prompt — 4-layer: identity + core skills + installed apps + matched skill
  4. Filter the unified tool namespace based on the matched skill’s allowed-tools globs
  5. Run the AgentEngine loop against the composed tool set
  6. Persist the conversation (JSONL or in-memory)

A connector is a remote MCP server a workspace connects to. The runtime orchestrates over remote MCP: it holds the server’s URL and a credential, and never downloads, verifies, or executes the server’s code. Supply-chain review belongs where the server is built and published, not in a process that also holds tenant credentials.

Three words sit around it and mean different things. A connection is the supervised link to one connector for one workspace and principal — it is what has a lifecycle, a credential, and a health state. A source is anything serving tools and resources behind the MCP boundary, which a connector’s connection is one of. And a platform app is a source the runtime hosts in-process rather than connects to, so it has no URL, no credential, and no connection to supervise.

Connectors are installed per workspace and tracked in workspace.json (not nimblebrain.json — entries placed there are stripped on load). They are discovered through the connectors catalog: a directory of ServerDetail files read into a single Browse list, installed point-and-click or via manage_connectors. See Connectors Catalog.

Each connection goes through these states:

starting ──► running ──► crashed ──► dead
│ ▲
└──► stopped ─────────┘
State Meaning
starting Connection being established
running Healthy and serving tools
not_authenticated Installed, no tokens yet — the user clicks Connect
pending_auth An OAuth flow is in flight
reauth_required Tokens were rejected — the user clicks Reconnect
crashed The connection dropped (will attempt recovery)
dead Repeatedly unreachable, not retrying on its own
stopped Manually stopped via API or CLI

Some connectors do not hold their own auth. A brokered connector gets both its credential and its hosted MCP session from a third party — Composio or Smithery — that an operator configures once, platform-wide. The session the broker returns is an ordinary remote source; the broker never sits in the tool call. A gateway is the simpler case: it brokers nothing, so one account key reaches every endpoint you publish from it. The runtime-native kinds, dcr and static, put no third party in the trust path at all. See Connector Providers.

NimbleBrain installs no connectors by default, and it still has tools on a fresh install. The kernel’s own capabilities — conversations, files, tasks, usage, skills, instructions, hooks, notifications, compose — are platform apps: in-process MCP servers the runtime hosts, reached over the same MCP boundary as a remote server. Nothing above the source can tell them apart, which is why the platform needs only one extension mechanism. Install connectors from the catalog to give the agent domain-specific tools.

{
"url": "https://mcp.example.com/mcp",
"serverName": "example"
}
Field Description
serverName The name the server registers under; tools reach the agent as <serverName>__<tool>
transport Transport class, auth, headers, and reconnection behavior
ui UI metadata: placement declarations

See Connector Configuration for the full field reference.

Skills are Markdown files with YAML frontmatter. They control what the agent knows and what tools it can use for a given message.

research.md
---
name: research
description: >
Deep research on a topic — search the web, gather sources, and synthesize a
report. Use when the user asks to research, investigate, or analyze a topic.
allowed-tools: websearch__search nb__search
metadata:
nimblebrain:
loading-strategy: dynamic
priority: 50
triggers:
- "research this"
- "deep dive"
---
You are a research agent. Search the web, gather multiple sources,
and synthesize findings into a structured report.
loading-strategy Priority Behavior
always 0–10 (core) or 11–99 Always composed into the system prompt — identity, voice, durable rules.
dynamic 11–99 Loaded on demand, via the activation signals below.

A dynamic skill enters context through any of:

  • Description — its name + description sit in a catalog; the model activates the relevant one (so the description must say what it does and when to use it).
  • Tool-affinity — auto-activates when a tool in the active set matches one of its tool-affinity globs.
  • Triggers — auto-activates on an exact phrase, for deterministic must-fire cases.
User: "research this topic and analyze the sources"
│
trigger "research this" ── match! ──► research skill activates

An activated skill injects its markdown body into the system prompt; allowed-tools (if set) scopes which tools it may call.

Skills are loaded from three locations (in order):

  1. Built-in — src/skills/builtin/ (shipped with the package)
  2. Global — ~/.nimblebrain/skills/
  3. Config — directories listed in the skillDirs config option

Tools come from two places: connectors and built-in system tools.

Every NimbleBrain instance gives the agent these nb__* tools:

Tool Description
nb__search Search installed tools (scope: "tools") or the connector catalog (scope: "catalog")
nb__status Platform status; scope: "connectors" for per-connector health, scope: "skills" for loaded skills, scope: "config" for model and limits
nb__read_resource Read an MCP resource by URI
nb__open_app Open an app on the user’s screen, optionally at a view inside it (by its resource URI)

Administration tools are app-only: nb__manage_connectors (install and disconnect connectors), nb__manage_workspaces (create workspaces and manage their members), and nb__manage_users are called by the settings pages and never appear in the agent’s tool list. The agent can find a connector in the catalog with nb__search, but installing it and creating a workspace happen in the UI.

App-platform capabilities are exposed through prefixed tool families: conversations__*, files__*, tasks__*, skills__*, usage__report, and compose__effective_context. (Workspace instructions are edited in the settings UI; the tool behind that UI is internal and not part of the agent’s surface.)

NimbleBrain uses tiered tool surfacing to keep the LLM’s tool list manageable:

Condition What the LLM sees
Total tools ≤ 30 All tools surfaced directly
Total tools > 30, no skill matched Only nb__* system tools. Others available via nb__search (scope: "tools").
Skill matched with allowed-tools Tools matching the globs + system tools. Others via nb__search.

The threshold is configurable via maxDirectTools (default: 30).

Conversations persist across messages so the agent remembers context.

Conversations are workspace-owned. Every turn is written to a per-conversation JSONL file under the workspace it ran in:

~/.nimblebrain/workspaces/<wsId>/conversations/<ownerId>/<convId>.jsonl

The owner (ownerId) is the authorization principal — only they can read or resume the conversation — while the workspace (wsId) is where it is stored. Persistence is automatic; there is nothing to configure.

Each conversation is a single append-only .jsonl event log:

{"id":"conv_abc123","ownerId":"user_abc123","createdAt":"2025-03-15T10:30:00Z", ...}
{"ts":"...","type":"user.message","content":[...]}
{"ts":"...","type":"run.start","runId":"run_1","model":"anthropic:claude-sonnet-4-6"}
{"ts":"...","type":"llm.response","runId":"run_1","content":[...],"usage":{...}}
{"ts":"...","type":"tool.done","runId":"run_1", ...}
{"ts":"...","type":"run.done","runId":"run_1","stopReason":"complete"}

Line 1 is the conversation header (ID, owner, workspace, creation time). Every later line is one event. Messages, the title, and token usage are derived from the events when the conversation is read, which is what lets a run replay after a crash and lets history be compacted.

To continue an existing conversation, open it from the Conversations sidebar in the web UI. Through the API, pass conversationId in the chat request body.

NimbleBrain is a full MCP Apps host, implementing the ext-apps specification. MCP Apps are interactive UI applications — built with HTML/JavaScript — that render directly inside the host as sandboxed iframes.

Unlike traditional web apps, MCP Apps:

  • Live inside the conversation — no tab-switching, context stays together
  • Call tools bidirectionally — apps invoke MCP tools and receive pushed results from the agent
  • Inherit the host’s theme — CSS variables are injected automatically
  • Run in a security sandbox — no access to the parent page, cookies, or other apps

Any connector that declares UI resources and placements becomes an MCP App. See MCP Apps for the full concept and MCP App Bridge for the implementation protocol.

Each workspace has a /mcp/<workspaceId> endpoint that turns it into a Streamable HTTP MCP server. External MCP clients — Claude, Claude Code, Open WebUI, or another NimbleBrain instance — connect to one workspace’s URL, which serves that workspace’s tools plus your identity tools (conversations, files, tasks) while you are a member, and nothing from any other workspace. Bare /mcp is refused.

This means NimbleBrain is both an MCP client (connecting to installed connectors) and an MCP server (exposing composed tools to external hosts). See MCP Endpoint Reference for configuration and usage.