Core Concepts
Composable orchestration
Section titled “Composable orchestration”NimbleBrain’s core value is composing multiple MCP servers into a single, orchestrated workspace. Rather than connecting to one MCP server at a time, NimbleBrain aggregates tools from every installed app into a unified namespace, layers prompts to give the agent awareness of all available capabilities, and filters tool access per-task through skills.
┌──────────────────────────────────────────────────┐│ ToolRegistry ││ ││ ┌─────────┐ ┌─────────┐ ┌──────────────────┐ ││ │ App 1 │ │ App 2 │ │ App 3 │ ││ │ (HTTP) │ │ (HTTP) │ │ (SSE) │ ││ │ 3 tools │ │ 5 tools │ │ 2 tools │ ││ └─────────┘ └─────────┘ └──────────────────┘ ││ ││ Unified namespace: 10 app tools + system tools │└──────────────────────────────────────────────────┘ │ │ ▼ ▼ Skill filtering Agent Engine (allowed-tools) (agentic loop)Three mechanisms make this composable:
| Mechanism | What it does |
|---|---|
| Tool aggregation | The ToolRegistry collects tools from every MCP source (remote HTTP/SSE servers, and the platform’s own in-process ones) and presents them as a flat namespace. The agent doesn’t know or care which server owns a tool. |
| Skill-scoped filtering | Each skill declares allowed-tools glob patterns. When a skill matches, only tools matching those patterns (plus system tools) are visible to the agent. This scopes the agent’s capabilities per task. |
| 4-layer prompt composition | System prompts are assembled from: (1) identity, (2) core context, (3) installed app metadata with trust scores, and (4) the matched skill’s prompt. Every layer adds awareness without the agent needing to discover it. |
Agent loop
Section titled “Agent loop”Every chat message triggers an agentic loop in the AgentEngine. The loop repeats until the model produces a response with no tool calls, or a limit is reached.
User message │ ▼┌─────────────────┐│ Call Claude │◄─────────────────┐│ (with tools) │ │└────────┬────────┘ │ │ │ ┌────┴────┐ │ │ Tool │ Yes │ │ calls? │──────► Execute tools │ │ │ in parallel ───┘ └────┬────┘ │ No ▼ Return responseEach pass through the loop is one iteration. Defaults:
| Setting | Default | Hard cap |
|---|---|---|
| Max iterations | 25 | 50 |
| Max input tokens | 500,000 | — |
| Max output tokens | the model’s catalog output limit (16,384 for a model the catalog lacks) | — |
models.default |
anthropic:claude-sonnet-4-6 |
— |
The engine stops for one of these reasons:
complete— the model responded without requesting any tool callsmax_iterations— the iteration limit was reachedmax_input_tokens— the run has an input-token cap (a task’s Max Input Tokens), and the next model call would pass it. Before each call the engine projects that call’s input as the larger of its estimate of the prompt and the previous call’s reported input, so the run ends within the cap unless the estimate undercounts. The cap counts cache reads, so it bounds tokens, not cost. Whatever the run produced before the stop is kept.maxInputTokensinnimblebrain.jsonis different: it bounds the context of one model call, not the run.spend_limit— the run names spend accounts (a task’s Token budget is two of them, one per cap), and what is left in one cannot pay for the next model call. Before each call the runtime projects the call’s input as formax_input_tokens, lowers the call’s output ceiling to what every account can pay for, and sets that much aside until the call returns; runs naming the same account share one balance, so they cannot together spend past it. The run stops when an account cannot cover the call’s input, or would leave it only a few hundred output tokens. Whatever the run produced before the stop is kept.cancelled— the run was stopped before it finished: the user pressed Stop, a task run was cancelled or timed out, or the runtime aborted it (for example on shutdown). The conversation keeps what the turn did before the stop.
All tool calls within a single iteration run in parallel via Promise.all().
Orchestration flow
Section titled “Orchestration flow”Before the engine loop starts, the Runtime orchestrates the composition described above:
- Resolve or create a conversation
- Match the user message to a skill (triggers, then keywords)
- Compose the system prompt — 4-layer: identity + core skills + installed apps + matched skill
- Filter the unified tool namespace based on the matched skill’s
allowed-toolsglobs - Run the
AgentEngineloop against the composed tool set - Persist the conversation (JSONL or in-memory)
Connectors
Section titled “Connectors”A connector is a remote MCP server a workspace connects to. The runtime orchestrates over remote MCP: it holds the server’s URL and a credential, and never downloads, verifies, or executes the server’s code. Supply-chain review belongs where the server is built and published, not in a process that also holds tenant credentials.
Three words sit around it and mean different things. A connection is the supervised link to one connector for one workspace and principal — it is what has a lifecycle, a credential, and a health state. A source is anything serving tools and resources behind the MCP boundary, which a connector’s connection is one of. And a platform app is a source the runtime hosts in-process rather than connects to, so it has no URL, no credential, and no connection to supervise.
Where connectors come from
Section titled “Where connectors come from”Connectors are installed per workspace and tracked in workspace.json (not nimblebrain.json — entries placed there are stripped on load). They are discovered through the connectors catalog: a directory of ServerDetail files read into a single Browse list, installed point-and-click or via manage_connectors. See Connectors Catalog.
Connection lifecycle
Section titled “Connection lifecycle”Each connection goes through these states:
starting ──► running ──► crashed ──► dead │ ▲ └──► stopped ─────────┘| State | Meaning |
|---|---|
starting |
Connection being established |
running |
Healthy and serving tools |
not_authenticated |
Installed, no tokens yet — the user clicks Connect |
pending_auth |
An OAuth flow is in flight |
reauth_required |
Tokens were rejected — the user clicks Reconnect |
crashed |
The connection dropped (will attempt recovery) |
dead |
Repeatedly unreachable, not retrying on its own |
stopped |
Manually stopped via API or CLI |
Brokered connectors
Section titled “Brokered connectors”Some connectors do not hold their own auth. A brokered connector gets both its
credential and its hosted MCP session from a third party — Composio or Smithery — that an
operator configures once, platform-wide. The session the broker returns is an ordinary
remote source; the broker never sits in the tool call. A gateway is the simpler case:
it brokers nothing, so one account key reaches every endpoint you publish from it. The
runtime-native kinds, dcr and static, put no third party in the trust path at all. See
Connector Providers.
Platform apps
Section titled “Platform apps”NimbleBrain installs no connectors by default, and it still has tools on a fresh install. The kernel’s own capabilities — conversations, files, tasks, usage, skills, instructions, hooks, notifications, compose — are platform apps: in-process MCP servers the runtime hosts, reached over the same MCP boundary as a remote server. Nothing above the source can tell them apart, which is why the platform needs only one extension mechanism. Install connectors from the catalog to give the agent domain-specific tools.
Connector options
Section titled “Connector options”{ "url": "https://mcp.example.com/mcp", "serverName": "example"}| Field | Description |
|---|---|
serverName |
The name the server registers under; tools reach the agent as <serverName>__<tool> |
transport |
Transport class, auth, headers, and reconnection behavior |
ui |
UI metadata: placement declarations |
See Connector Configuration for the full field reference.
Skills
Section titled “Skills”Skills are Markdown files with YAML frontmatter. They control what the agent knows and what tools it can use for a given message.
Skill file format
Section titled “Skill file format”---name: researchdescription: > Deep research on a topic — search the web, gather sources, and synthesize a report. Use when the user asks to research, investigate, or analyze a topic.allowed-tools: websearch__search nb__searchmetadata: nimblebrain: loading-strategy: dynamic priority: 50 triggers: - "research this" - "deep dive"---
You are a research agent. Search the web, gather multiple sources,and synthesize findings into a structured report.How a skill loads — loading-strategy
Section titled “How a skill loads — loading-strategy”loading-strategy |
Priority | Behavior |
|---|---|---|
always |
0–10 (core) or 11–99 | Always composed into the system prompt — identity, voice, durable rules. |
dynamic |
11–99 | Loaded on demand, via the activation signals below. |
How dynamic skills activate
Section titled “How dynamic skills activate”A dynamic skill enters context through any of:
- Description — its name + description sit in a catalog; the model activates the relevant one (so the description must say what it does and when to use it).
- Tool-affinity — auto-activates when a tool in the active set matches one of its
tool-affinityglobs. - Triggers — auto-activates on an exact phrase, for deterministic must-fire cases.
User: "research this topic and analyze the sources" │ trigger "research this" ── match! ──► research skill activatesAn activated skill injects its markdown body into the system prompt; allowed-tools (if set) scopes which tools it may call.
Skill directories
Section titled “Skill directories”Skills are loaded from three locations (in order):
- Built-in —
src/skills/builtin/(shipped with the package) - Global —
~/.nimblebrain/skills/ - Config — directories listed in the
skillDirsconfig option
Tools come from two places: connectors and built-in system tools.
System tools
Section titled “System tools”Every NimbleBrain instance gives the agent these nb__* tools:
| Tool | Description |
|---|---|
nb__search |
Search installed tools (scope: "tools") or the connector catalog (scope: "catalog") |
nb__status |
Platform status; scope: "connectors" for per-connector health, scope: "skills" for loaded skills, scope: "config" for model and limits |
nb__read_resource |
Read an MCP resource by URI |
nb__open_app |
Open an app on the user’s screen, optionally at a view inside it (by its resource URI) |
Administration tools are app-only: nb__manage_connectors (install and disconnect connectors), nb__manage_workspaces (create workspaces and manage their members), and nb__manage_users are called by the settings pages and never appear in the agent’s tool list. The agent can find a connector in the catalog with nb__search, but installing it and creating a workspace happen in the UI.
App-platform capabilities are exposed through prefixed tool families: conversations__*, files__*, tasks__*, skills__*, usage__report, and compose__effective_context. (Workspace instructions are edited in the settings UI; the tool behind that UI is internal and not part of the agent’s surface.)
Tool surfacing strategy
Section titled “Tool surfacing strategy”NimbleBrain uses tiered tool surfacing to keep the LLM’s tool list manageable:
| Condition | What the LLM sees |
|---|---|
| Total tools ≤ 30 | All tools surfaced directly |
| Total tools > 30, no skill matched | Only nb__* system tools. Others available via nb__search (scope: "tools"). |
Skill matched with allowed-tools |
Tools matching the globs + system tools. Others via nb__search. |
The threshold is configurable via maxDirectTools (default: 30).
Conversations
Section titled “Conversations”Conversations persist across messages so the agent remembers context.
Storage
Section titled “Storage”Conversations are workspace-owned. Every turn is written to a per-conversation JSONL file under the workspace it ran in:
~/.nimblebrain/workspaces/<wsId>/conversations/<ownerId>/<convId>.jsonlThe owner (ownerId) is the authorization principal — only they can read or resume the conversation — while the workspace (wsId) is where it is stored. Persistence is automatic; there is nothing to configure.
JSONL format
Section titled “JSONL format”Each conversation is a single append-only .jsonl event log:
{"id":"conv_abc123","ownerId":"user_abc123","createdAt":"2025-03-15T10:30:00Z", ...}{"ts":"...","type":"user.message","content":[...]}{"ts":"...","type":"run.start","runId":"run_1","model":"anthropic:claude-sonnet-4-6"}{"ts":"...","type":"llm.response","runId":"run_1","content":[...],"usage":{...}}{"ts":"...","type":"tool.done","runId":"run_1", ...}{"ts":"...","type":"run.done","runId":"run_1","stopReason":"complete"}Line 1 is the conversation header (ID, owner, workspace, creation time). Every later line is one event. Messages, the title, and token usage are derived from the events when the conversation is read, which is what lets a run replay after a crash and lets history be compacted.
Resuming conversations
Section titled “Resuming conversations”To continue an existing conversation, open it from the Conversations sidebar in the web UI. Through the API, pass conversationId in the chat request body.
MCP Apps
Section titled “MCP Apps”NimbleBrain is a full MCP Apps host, implementing the ext-apps specification. MCP Apps are interactive UI applications — built with HTML/JavaScript — that render directly inside the host as sandboxed iframes.
Unlike traditional web apps, MCP Apps:
- Live inside the conversation — no tab-switching, context stays together
- Call tools bidirectionally — apps invoke MCP tools and receive pushed results from the agent
- Inherit the host’s theme — CSS variables are injected automatically
- Run in a security sandbox — no access to the parent page, cookies, or other apps
Any connector that declares UI resources and placements becomes an MCP App. See MCP Apps for the full concept and MCP App Bridge for the implementation protocol.
Platform as MCP server
Section titled “Platform as MCP server”Each workspace has a /mcp/<workspaceId> endpoint that turns it into a Streamable HTTP MCP server. External MCP clients — Claude, Claude Code, Open WebUI, or another NimbleBrain instance — connect to one workspace’s URL, which serves that workspace’s tools plus your identity tools (conversations, files, tasks) while you are a member, and nothing from any other workspace. Bare /mcp is refused.
This means NimbleBrain is both an MCP client (connecting to installed connectors) and an MCP server (exposing composed tools to external hosts). See MCP Endpoint Reference for configuration and usage.