← Browse

@escoffier-labs/brigade-2

A

When an agent says tests passed, you get a file with the real exit code, a code graph of what changed (brigade code), and an evidence log you can search later (brigade evidence).

instructionscodex

Install

agr install @escoffier-labs/brigade-2 --target codex

Writes 1 file into AGENTS.md, pinned to git-6f50c443.

  • AGENTS.md

Document

AGENTS.md

Orientation for coding agents working on GraphTrail.

GraphTrail is a local code-graph sidecar. It parses a repository with tree-sitter in a single pass per file, extracts symbols, imports, and call edges into a small SQLite graph under .graphtrail/, and answers structural questions (search, callers, callees, impact, context, stats) plus freshness checks (doctor), a dry-run evaluate, edge lineage (explain), graph export, and an opt-in foreground watch over two surfaces: a CLI (graphtrail) and an MCP server (graphtrail-mcp). The default build makes no network calls and starts no daemon. Languages supported: Python, TypeScript/JavaScript, Rust, Go.

MCP query connections always use SQLITE_OPEN_READ_ONLY. refresh: true starts an incremental graph-index write and waits up to 10 seconds before opening the query read-only. If the refresh fails or times out, the query proceeds and appends a refresh_error note to its text result. A timed-out worker may finish concurrently with that read-only query. Without refresh, query tools do not write the graph.

Build and test

cargo build --release        # binaries land in target/release/
cargo test --all-features

CI gate

CI runs the same checks it expects from you. Run them locally before pushing:

cargo fmt --check
cargo clippy --all-targets --all-features -- -D warnings
cargo test --all-features
cargo build --release

Brigade work loop

This repository is Brigade-wired. At session start, invoke using-skillet to select the applicable workflow, then invoke brigade-work before running raw Brigade commands. Read the work brief before editing:

brigade work brief --target .

Run checks through Brigade so the exit code is recorded, then capture the outcome against the skill or card that guided the change:

brigade work verify run --target . --command "cargo test --all-features"
brigade outcome capture taste --run-id latest --kind skill

Replace taste with the skill or card used for that verification. After substantial work, write durable findings in the standard Memory Handoff format under .claude/memory-handoffs/, then run brigade handoff lint before finishing.

Module map

The code is split into focused modules:

  • model (src/model.rs): shared types.
  • evaluate (src/evaluate.rs): dry-run extraction with zero database writes.
  • watch (src/watch.rs, feature watch, on by default): foreground debounced sync-on-change.
  • extractors (src/extractors/): per-language tree-sitter providers plus shared traversal in common.rs. Each language is a provider behind the LangSpec trait.
  • store (src/store/): database access, locking, metadata, schema upgrades, repository policy, incremental sync, persisted pending calls, edge resolution, and edge lineage (explain).
  • query (src/query/): symbol search, graph traversal, context packs, stats, freshness checks, graph diffs, structural health, and affected-test attribution.
  • mcp (src/mcp.rs): JSON-RPC handling plus the MCP tool registry, argument policy, and dispatch.
  • adapters (src/adapters/): optional Code Search and MiseLedger integrations behind cargo features.
  • cli (src/cli.rs): a thin command-line interface.
  • Binaries: src/main.rs (the graphtrail CLI) and src/bin/graphtrail-mcp.rs (the graphtrail-mcp MCP server).

MCP smoke test

Build, then pipe newline-delimited JSON-RPC into the server over stdio. It speaks JSON-RPC 2.0 and exposes fourteen tools (search, callers, callees, impact, context, stats, doctor, file_neighbors, dead_code, cycles, affected, explain, repos, diff).

cargo run -- init .
cargo run -- sync .
cargo run -- --db .graphtrail/graphtrail.db stats --json
cargo build --release
printf '%s\n%s\n' \
  '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{}}' \
  '{"jsonrpc":"2.0","id":2,"method":"tools/list"}' \
  | ./target/release/graphtrail-mcp --db .graphtrail/graphtrail.db

Conventions

  • Keep the default build network-free and small. Network or cross-tool integrations go behind an optional cargo feature (see codesearch and miseledger), never in the default binary.
  • MCP query connections must stay read-only. The single sanctioned graph write starts when a supported query receives refresh: true. Preserve the 10-second wait, fail-open refresh_error note, and possible overlap between a timed-out worker and its query.
  • Schema, JSON output shapes, and MCP tool contracts are stable contracts; breaking changes need a conversation first.
  • Each language extractor owns an EXTRACTOR_FINGERPRINT constant. Bump that language's fingerprint whenever the extractor can produce different symbols, imports, calls, symbol ids, signatures, containers, body hashes, language labels, or filtering behavior for the same file content. Do not bump unrelated language fingerprints. One exception: a change to the SHARED symbol-id derivation in extractors/common.rs affects every language at once and must ship as a schema migration that rewrites ids in place (see the v7 rewrite in store/schema.rs), not as four fingerprint bumps.
  • No personal details, hostnames, IPs, account IDs, or live auth profiles in code, tests, or fixtures.
  • Conventional commits only. No AI co-authorship trailers.

See CONTRIBUTING.md for what lands easily and README.md for the full design rationale.

Repository README

Describes escoffier-labs/brigade as a whole, which may contain artifacts other than this one. Where this artifact had no useful description of its own, its summary was taken from here.

The loop

Every piece of work Brigade touches runs the same circuit:

  1. Intent and acceptance criteria go in.
  2. Prior evidence and code impact attach to the brief.
  3. Replaceable workers execute, bounded.
  4. Verification runs with a real exit code.
  5. The receipt, graph delta, and outcome land where the next run starts.

Work enters as intent and leaves as evidence.

Install

Default path: give the job to an agent. Paste something like this into Claude Code, Codex, Cursor, OpenClaw, or any agent with shell access:

Read https://github.com/escoffier-labs/brigade and follow AGENTS.md.
Install brigade-cli with pipx if missing. Run operator quickstart with
--dry-run first and show the plan. Then apply, run brigade operator doctor,
and report ready: yes. Keep existing memory layout. Do not touch remotes,
do not commit, stop before anything destructive.

The agent installs, runs brigade setup, wires harnesses, and leaves files on disk. You review when the gate is ambiguous or risky. You do not have to type the install path yourself.

If you prefer a shell (same commands the agent would run):

pipx install brigade-cli
pipx ensurepath           # then open a new shell so `brigade` is on PATH
brigade setup             # install the verified native engines
brigade operator quickstart --target ./my-repo --harnesses codex

Brigade prints a one-line notice when a new release is out. It checks at most once a day through an anonymous request. Set BRIGADE_NO_UPDATE_CHECK=1 to disable it. Details are in docs/update-channels.md.

Stable pinners may deliberately install an exact release with pipx install brigade-cli==X.Y.Z or refresh through brigade update --channel stable. Channel ownership, beta rules, and when to use brigade update are in docs/update-channels.md.

brigade operator doctor --target ./my-repo prints ready: yes when the wiring is healthy. The default footprint is small: AGENTS.md, SAFETY_RULES.md, a handoff template, and .brigade/ state. Add --dry-run to preview anything before it writes. Nothing leaves your machine.

Per-OS setup (apt, Homebrew, Scoop, PowerShell), workspace depth, and multi-harness installs: install guide, QUICKSTART.md, first 10 minutes. Homegrown setup already? brigade operator adopt plan.

Auditable agents: every action leaves a receipt

An agent that reports "tests pass, nothing else changed" is making a claim. Brigade turns the claim into a record. Run any check through Brigade and it writes a receipt: the command, the real exit code, the graph delta, the git state.

brigade work verify run --target . --command "pytest -q" --capture brigade-work
// .brigade/work/verify-runs/20260722-033355-work-verify-298ff5/receipt.json (abridged)
{
  "run_id": "20260722-033355-work-verify-298ff5",
  "status": "completed",
  "duration_seconds": 15.18,
  "commands": { "check": { "argv": ["pytest", "-q"], "returncode": 0 } },
  "code_graph_delta": { "changed_symbol_count": 0 },
  "git": { "branch": "main", "dirty_files": 0 }
}

Receipts land in the evidence log (brigade evidence), a Go engine installed by brigade setup (historically shipped as MiseLedger). Every consequential action elsewhere in Brigade (a memory write, a skill promotion, a sync) is logged the same way. brigade evidence search answers "what ran, when, and what did it change" from files, weeks later. Cross-model dispatches through brigade run carry the same paper trail. When someone asks what your agents did this week, the answer comes from receipts you can grep, not from scrollback. Capability page.

Code graph, built in

The code_graph_delta line in that receipt comes from the code graph (brigade code), a Rust engine installed by brigade setup (historically shipped as GraphTrail, digest-verified, no toolchain required). It indexes your repo once and keeps up incrementally. On this repository a sync pass over 652 files and 10,405 symbols reports in well under a second. A receipt names the exact symbols a change touched, and your agents stop grepping and start asking structural questions:

$ brigade code impact _write_receipt
test_runbook_closeout_imports_failed_steps  --calls@31--> _write_receipt
test_runbook_closeout_without_import_flag   --calls@45--> _write_receipt

$ brigade code context "verify receipts"
## Entry Points
- function `_verify_receipts` at src/brigade/work_cmd/verification.py:185-192
- function `_iter_verify_receipts` at src/brigade/workflow_cmd.py:142-167

The graph feeds the rest of Brigade instead of sitting idle: brigade run prepends a context pack when a graph exists, so a dispatched model starts from callers and blast radius instead of a cold grep, and an MCP server ships alongside so any harness can query the graph directly. Capability page.

One MCP and tool catalog, synced into every tool

Every agent tool reads its MCP servers from a different file in a different shape. The same servers wired across Claude Code, Cursor, Codex, VS Code, OpenCode, and Antigravity means hand-editing six configs and keeping them in sync forever. Brigade keeps one canonical catalog and merges it into each tool's native config for you.

brigade mcp init                  # scaffold .brigade/mcp.json
brigade mcp add --name github --command npx \
  --args "-y @modelcontextprotocol/server-github" \
  --env GITHUB_AUTH_ENV=ref:BRIGADE_GITHUB_AUTH_ENV
brigade mcp sync                  # dry-run: show the diff for every tool
brigade mcp sync --write          # merge into each tool's config
brigade mcp verify                # check initialize + tools/list at runtime

Run brigade mcp sync and you get the per-tool plan, server by server, before a single file changes:

brigade mcp sync (dry-run): ~/my-repo
claude       github               missing        -> create
claude       sentry               missing        -> create
cursor       github               missing        -> create
cursor       sentry               missing        -> create
codex        github               missing        -> create
codex        sentry               missing        -> create
vscode       github               missing        -> create
vscode       sentry               missing        -> create
opencode     github               missing        -> create
opencode     sentry               missing        -> create
ToolFile it writes
Claude Code.mcp.json
Cursor.cursor/mcp.json
Codex CLI.codex/config.toml (merged surgically, other tables preserved)
Grok CLI.grok/config.toml (same TOML shape as Codex)
VS Code.vscode/mcp.json (secrets become inputs[])
OpenCodeopencode.json
Antigravity~/.gemini/config/mcp_config.json (user-scoped)

Dry-run by default. Merges by server key, so servers you added by hand are never touched. Secrets are written as ${VAR} references, never inlined. Ownership lives in a gitignored sidecar, so a fresh clone re-syncs without conflict. Tools and skills get the same treatment via brigade tools sync: one reviewed catalog, projected into each harness's native format. Full merge rules: docs/mcp-sync.md. Evaluating options first? The comparison page.

Shared memory and verified learning

Writer harnesses leave handoff notes as they work. Brigade lints, guards, and classifies each one. Safe, targeted notes file themselves into durable memory. The ambiguous few wait for your review. Every consequential action is logged to a plain file you can grep, diff, and prune.

Memory stays two layers deep: knowledge cards hold the detail, MEMORY.md stays a slim one-line-per-card index that loads every session. brigade memory care scan flags stale or contradictory cards instead of letting them rot, and brigade evidence search plus exported briefs mean the next session starts where this one stopped.

Filing notes is the first loop. The second loop earns trust: Brigade promotes a learned skill only when a real signal proves it helped, and rolls it back the moment a signal says it broke. The model never grades its own work.

$ brigade outcome rank
- brigade-work      score=0.692  helped=675  hurt=260
- ultra-work-scout  score=0.563  helped=50   hurt=24
- memory-handoff    score=0.490  helped=8    hurt=2
  • brigade outcome capture records a verify run's real exit code against the skill that produced it.
  • brigade outcome score ranks by a Wilson lower bound, so two lucky passes never outrank twenty vetted runs.
  • brigade outcome reconcile is the gate: dry-run by default, --apply installs a skill that earned it or rolls a regressed one back.
  • brigade outcome explain prints the full signal trail behind any decision, so every promotion is as auditable as the runs that earned it.

The ledger is plain JSON and markdown under memory/outcome/, tracked in git, readable without Brigade. brigade init wires a brigade-work skill into each harness so agents run this loop without being told. With Claude Code it also installs project-scoped hooks that redirect raw test commands through verify runs.

Optional stations

Built in (via brigade setup):

SurfaceCommandsNotes
Code graphbrigade code …Callers, impact, context. Formerly GraphTrail.
Evidence logbrigade evidence …Searchable ledger of runs and imports. Formerly MiseLedger.
Content Guardbrigade guard / brigade scrubSecrets and private detail before publish.

Optional stations (add when you need them). Core works with none installed. brigade status health-checks whatever is present.

StationInstallRole
Agent Pantrybrigade add pantryEncrypted browser-session and secret sync across machines
Token Glacebrigade add tokensCompact noisy tool output before it burns context
Skilletoptional rosterPortable skills that reconcile can promote or roll back
Bootstrap Doctorbrigade add bootstrap-doctorAudit OpenClaw bootstrap files (SOUL.md, TOOLS.md, AGENTS.md, IDENTITY.md, MEMORY.md, and the rest of the set) and trim oversize detail into cards
Notificationsbrigade add notificationsOptional agent-notify binary for Discord, Telegram, or Signal. Status and setup planning only until you wire hooks or pass an explicit --send

Upgrading from standalone GraphTrail or MiseLedger installs? brigade setup replaces both. The old brigade add graphtrail / add evidence paths remain as compatibility shims. Engine binaries and some paths still use the historical names. The operator surface is brigade code and brigade evidence. Details: wiring guide, station contract.

Beyond the daily loop, the same review-and-receipt pattern covers cross-model runs (brigade run dispatches one bounded task across your roster), security scans, friction mining, research reports, and fleet health. All of it stays behind brigade extras on until you ask. The full tour: docs/overview.md.

Why not something else?

  • mem0, Letta, agentmemory, and friends are memory layers for apps you are building, usually behind an API or a server. Brigade is for the agent CLIs you already run, and it is file-first: your memory is markdown in your repo, reviewable in git, readable without Brigade.
  • add-mcp, chezmoi, and config-sync scripts move MCP or dotfiles around, but they sync one thing with no review gate and no receipt. Brigade keeps MCP servers, tools, skills, and memory in one canonical source, shows the per-tool diff before any write, and leaves a receipt you can roll back.
  • Native harness memory is a per-tool silo. It does not cross harnesses, and it writes without review. Brigade gives every tool one shared format and one canonical owner, with a review gate in between.
  • Already running Hermes, or any self-improving agent? Keep it. Brigade is the verification layer on top: it promotes a skill only when a real signal confirms it, keeps every learned skill as portable markdown in your git, and runs one loop across your whole fleet.
  • A plain CLAUDE.md / AGENTS.md works great until it bloats past the context budget and goes stale. Brigade keeps bootstrap files slim, moves detail into indexed cards, and flags staleness instead of trusting last month's facts forever.
  • A daemon or hosted service would be simpler to demo and worse to trust. Brigade writes local files when you run a command, and that is all it does.
Across harnessesMCP, tools, and memory in one sourceReview gate + receiptsLocal files, no daemon
Brigadeyesyesyesyes
mem0 / Letta / agentmemoryper-SDKmemory onlynousually hosted or a server
add-mcp / chezmoi / config-syncpartialMCP or dotfiles onlynoyes
Native harness memorynomemory onlynoyes

What Brigade is not

Brigade is not a hosted memory service, a daemon, or an automatic release bot. It does not run in the background or install schedulers (one scoped exception: brigade tools runtime start launches a local runtime process, only when you start it, until you stop it). It does not push to GitHub, publish packages, save every note automatically, or skip review for ambiguous, risky, or failed notes. brigade work brief and related status surfaces may report notification readiness or suggest installing the notifications station, but Brigade never sends a message unless the operator uses an explicit send action such as brigade pantry expiry-alert --send. That pause is the point: agent memory should be useful, not noisy.

And it is not the other projects that share the name. This Brigade is the AI-agent operator CLI from escoffier-labs/brigade, installed with pipx install brigade-cli. It is not the CNCF/Microsoft Brigade for Kubernetes event scripting (archived 2022), the Spinabot Brigade agent crew, or the 2017 brigade Python package that became Nornir.

Why I built this

I run an always-on OpenClaw agent next to daily Codex and Claude Code sessions. Every one of those tools wakes up empty, and whatever a session learned scattered across tool-specific folders and died there. Two incidents shaped the design: a "dreaming" job that promoted raw session fragments straight into memory bloated MEMORY.md past the bootstrap budget, so every session started truncated and nobody noticed for weeks. And 195 handoff notes sat unread across 35 repos because an ingester had a hardcoded allowlist and nothing warned about the gap. Silence is the failure mode. Every part of Brigade that lints, warns, or writes a receipt exists because something once failed in silence. The full production stack, now 482 cards across daily multi-agent work, is documented in the Cookbook.

Names (public vs historical)

What you type / sayHistorical nameNotes
Code graph · brigade codeGraphTrailBuilt in via brigade setup
Evidence log · brigade evidenceMiseLedgerBuilt in via brigade setup
Content Guard · brigade guard / scrubcontent-guardEmbedded
Bootstrap DoctorsameFull OpenClaw bootstrap set: SOUL, TOOLS, AGENTS, IDENTITY, MEMORY, and related session-start files

Kitchen language (brigade de cuisine, mise en place, station nicknames) is brand and deep docs. Commands and product surfaces stay plain: receipt, code graph, evidence log, sync, handoff.

Who runs what: agents install, set up, verify, and write handoffs. Humans set policy and review gates when something is ambiguous or risky. The CLI is the control plane the fleet drives, not a tool you are expected to operate by hand all day.

Harnesses

Nineteen harnesses get handoff inboxes and ingest coverage, from Codex, Claude Code, and Cursor to Goose, Aider, and OpenHands. Most also get projected tools and skills in their native format. The per-harness matrix is in the technical guide.

Questions

@brigadeclaw answers questions about Brigade on X. Every sentence it posts cites the pinned documentation it came from, and it stays silent rather than guessing. How it works, including the gates and the refusals: brigade.tools/brigadeclaw.

Docs

License

MIT. See LICENSE.

Project identity: GitHub escoffier-labs/brigade, website brigade.tools, PyPI brigade-cli, command brigade. The product name comes from a kitchen line (brigade de cuisine): coordinated stations, prep before service. You do not need the kitchen glossary to install or run it. Set up rules, memory, tools, and receipts before the session gets expensive.

It is early-stage and moving fast. If you hit a broken workflow, a confusing command, or a setup issue, open an issue and I will get it fixed.

Trustgrade A

  • passBody integrity

    Whether the stored document is plausibly the kind of file the artifact declares, rather than something fetched by mistake.

  • passType matchnot applicable to this artifact type

    Whether the artifact is really the kind of thing its metadata claims it is.

  • passFreshness

    How long since the source repository was last pushed to.

  • passPrompt injection

    Scans the artifact's own text for instructions aimed at your agent rather than at you.

  • passLicense

    Whether the source repository declares an SPDX license permissive enough to redistribute.

How the grade is calculated

Each check contributes 0 points when it passes, 1 when it warns, and 2 when it fails. The total maps to a letter:

  • Aevery check passed
  • Bone warning
  • Ctwo warnings
  • Dprompt injection or body integrity failed, or three warnings
  • Fone of those failed, and something else is wrong

These are automated hygiene checks, not a security audit, and not a dependency or vulnerability scan. A grade of A means nothing was flagged — not that the artifact is safe.

Versions

  • git-6f50c443ad372026-08-04