FlowCodex is an AI coding assistant that lives in your shell. A Director runs a fleet of subagents — each isolated in its own git worktree, budget and context — while append-only sessions, add-only memory and a permission policy that asks only for what it cannot undo keep the whole thing accountable.
One file, its own runtime — no Node, no Bun. Installs to ~/.local/bin,
linked as fcx; SHA-verified on install.
Build from source ↗
The OpenTUI transcript: one row per tool call, a swarm strip for live subagents, and one summary line per turn — duration, files, tests, cost.
These are not slogans on a wall. Each one is a structural property of the codebase, with a contract test or a lint rule holding it in place.
Every release is in the changelog. ADRs record where the build diverged from the plan — and keep the record. Docs are checked against the source, not against each other.
Permission is about what cannot be undone or leaves the machine — not what
looks scary. Measured on real journals: about five prompts a month in the default
trusted mode.
Sessions are an append-only JSONL event log. Memory consolidation is add-only by construction — supersede and contradict edges are queued for your approval, never applied silently.
REPL, TUI, HTTP/WS server and ACP drive the same engine through one Surface
contract, rendering one event stream. Nothing in the runtime knows what a terminal is.
Fourteen subsystems, each with its own package, its own docs page and its own tests.
Streaming responses with interleaved reasoning, explicit [continue]/[done]
turn markers instead of language heuristics, loop guards, and a graduated compaction
pipeline that keeps months-long sessions inside the context window.
Fully custom wire adapters for Anthropic, OpenAI, Google and any OpenAI-compatible
endpoint — plus OAuth for Claude, Codex and GitHub Copilot. Multi-key auth in an
encrypted auth.json, per-provider health with cooldowns, and a
cross-provider fallback chain. No model list is hardcoded; models.dev is the sole catalog.
read write edit bash grep
glob todo webfetch patch test
and more. Mutating tools run through a transactional mutation pipeline; external content
is fenced so tool output cannot impersonate instructions.
delegate, spawn_subagent, fleet_status,
merge_worktree. Subagents run in isolated git worktrees under their own
budget and context, with a Brain risk governor and mid-run steering —
/steer, /btw side questions, a mailbox, Esc-to-steer.
Three SQLite scopes (user, project, session) with FTS5/BM25 search, a self-populating memory graph, and anchors that record why they stopped matching when the code under them changes. Add-only by construction.
Per-project symbol index with cross-references and an on-demand CodeMap graph. Rebuilt in a candidate file and atomically swapped — a failed rebuild never leaves you with half an index.
Append-only JSONL with resume, fork, rewind, pins, export and import. A diagnostic
manifest (/session diagnose) reports tool metrics, cost breakdown,
compaction timeline and permission decisions after the fact.
A plugin ships skills, commands, agents, hooks, MCP servers and tools as one
package (ADR-028) — installed from a marketplace, a repository or a directory
with /plugin install, reviewed before anything runs, kept in a
content-addressed store, and rolled back on demand. Claude Code and Codex
packages are read as published.
Skills as on-demand prompt extensions, discovered across project, user and bundled layers. An MCP client and server with an LSP plugin — 32 built-in plugins load behind a capability-checked, trust-tiered host with atomic generations.
An Agent Client Protocol server over stdio or loopback socket, a client that spawns trusted local agents, parallel ensembles, crash-resumable benchmark journals, and a threshold-signed registry with trust-root rotation.
One /goal runs the agent unattended until an independent evaluator —
or a --check command you write — says the objective is met. Esc pauses
it, a budget caps the whole fleet, and it lives in the session log like everything
else (ADR-027).
A bare flowcodex opens the OpenTUI alternate-screen renderer: pure-TypeScript
markdown with highlighted diffs, semantic tool cards, mouse capture with click-to-caret,
an in-app scrollable transcript, and a swarm grid of live-agent tiles. The same binary
links as fcx; --repl keeps the classic REPL.
flowcodex serve exposes HTTP/WS with OpenAPI and auth middleware;
@flowcodex/sdk is the zero-dependency client. flowcodex attach
connects a TUI to a running server. Structured output against any JSON Schema.
Risk-classified shell verdicts, snapshot-before-undo, an optional kernel sandbox, secret and prompt-injection pre-commit scanning, dependency audit — and subagents that hand risky commands to you instead of running them.
Everything below runs from the single binary — no Node, no Bun, no global packages.
curl -fsSL https://github.com/Msvnc0/flowcodex/releases/latest/download/install.sh | sh
The installer verifies the SHA-256 sums, links the binary as both
flowcodex and fcx, and writes the PATH entry in your shell's
own dialect. Windows uses install.ps1 the same way.
flowcodex
Opens the TUI; type a task and the Director routes it. --repl starts the
classic REPL, flowcodex "<prompt>" starts with an initial prompt, and
flowcodex serve is the headless HTTP/WS server you can
attach a TUI to.
/connect
A catalog-driven picker walks through Anthropic, OpenAI, Google, Copilot or any
OpenAI-compatible endpoint — including local models. Keys live in an encrypted
auth.json, multi-key, with automatic failover between providers.
@flowcodex/core depends on nothing. The CLI is the only composition root.
Package boundaries and core layering are enforced by pnpm lint:arch —
a violation fails the build, not a review comment.
Faithful rendering of the diagram in docs/architecture.md — surfaces subscribe to one event flow; the CLI is the composition root.
| Check | What it catches |
|---|---|
| pnpm lint:arch | Package boundary violations (static and dynamic imports), core layer violations, bidirectional coupling, import cycles, direct tool registration outside the allowlist, unowned child processes |
| pnpm check:contracts | Workspace version and export contracts, in source and in the packed tarballs |
| scripts/*-truth.test.ts | Settings rows, persistence writers, dead contracts and version reporting — asserted against the source rather than a doc |
| pnpm typecheck | strict, noUncheckedIndexedAccess, exactOptionalPropertyTypes across every package |
Every subagent is cut on what its work may touch, runs in its own git worktree, and carries its own budget and context window. The Brain governor escalates on budget or failure, asks a human when it must, and enforces a hard cost cap.
Solid lines: spawn and await. Dashed: the Brain's oversight and the mailbox that lets a running agent be steered without aborting it.
| Role | Mutates | Runs | Reads | Use when |
|---|---|---|---|---|
| explore | — | — | ✓ | find or read anything, including the web |
| review | — | — | — | a verdict on a diff, no test run |
| verify | — | — | — | tests and typecheck actually run |
| build | ✓ | — | — | code has to change |
| debug | ✓ | ✓ | — | something must run to know what to change |
| general | ✓ | ✓ | ✓ | nothing above fits — deliberately last |
general has every tool — which is exactly why the router treats it as the
most expensive answer. Agent files may narrow a role, never widen it (ADR-024).
And one /goal: an autonomous loop where done is decided by an independent
evaluator (or a --check command), never by the agent itself. Esc pauses it;
a budget caps the whole fleet.
Most agent memory is a diary. FlowCodex memory is a graph with anchors in your code: when the code under an anchor changes, the drift is recorded with its reason — content changed, file gone, symbol gone — and it survives a restart.
User (~/.flowcodex/memory.db), per-project, and per-session — routed by
one MemoryService, searched with FTS5/BM25, retrieved identifier-aware
and scoped by audience. Subagents read through their own ledger.
The consolidator and the tools never delete or rewrite. Supersede, contradict and
near-duplicate edges go to a pending queue for /memory approve.
Hard deletes need an explicit command from you.
The system prompt is built once per cache epoch and never edited — later changes
arrive once as a [STATE] note. Four cache_control
breakpoints, placed so the next request reads the previous one's cache instead of
re-billing your history.
One month of the maintainer's journals — 1,120 leader shell calls — held no destructive command at all. The policy that came out of that data asks for roughly five prompts a month, and it can tell you exactly why.
A classifier classes each command before it runs:
What a snapshot can bring back — git reset --hard,
an in-project rm -r — is undoable: it runs behind a snapshot
taken right before it. Pushes, publishes and credential reads are risk:
they ask.
trusted by default--read-only is a lock, not a mode. Subagents never
ask — they hand risk to the leader, which asks you once. A repository's own
permission grants apply only after you approve them, bound to their content.
An optional kernel sandbox runs shell with project-only writes, no network and
credential paths unreadable in every mode. External content is fenced by an authority
model so a web page cannot impersonate your instructions. And
features.approver: 'model' hands the questions to a judge model — whose
allow is never a human confirmation.
FlowCodex releases frequently and documents every one of them. The plan below comes from the changelog and the decision records — including the parts that are deliberately still closed.
Every compaction pass records the preserved-tail boundary in the session event —
message count plus a role/sha fingerprint of the first kept message — and resume
restores the tail from it. Pass digests persist into /session report,
and compaction advice now appears in the REPL the moment it fires.
A plugin ships skills, commands, agents, hooks, MCP servers and FlowCodex tools as
one directory, installed by digest with /plugin install, its executable
set approved as a unit and rolled back on demand. Claude Code, Codex and Agent
Plugins packages are read as published — one dialect per package, never merged.
A bare flowcodex opens the TUI; the same binary links as
fcx. The installer now writes PATH entries in the invoking shell's own
dialect — fish drop-ins, marked rc blocks — and installs to
~/.local/bin, leaving program and user data as two separate removals.
Every LLM response now carries a fingerprint of its cache prefix; /session report
shows hit rate, re-billed tokens and breaks by cause. On Anthropic, two of four
cache_control slots moved onto the conversation so each turn reads the
previous one's cache; OpenAI-compatible providers get prompt_cache_key.
/goal, a model judge, hard spend capsThree near-duplicate autonomy commands collapsed into one loop with an independent evaluator (ADR-027). A judge model can answer permission questions from what you typed — never past the ceiling, never for protected paths. Budgets cap the fleet, not just the leader.
Each release ships a single binary that carries its own runtime — Linux, macOS
(x64/ARM64) and Windows. The installer checks it against the release's
SHA256SUMS; flowcodex update replaces it the same way.
The packages are built and tested; npm install -g flowcodex goes live
once the repository's PUBLISH_NPM variable is set to true.
The release workflow publishes them automatically from then on.
The threshold-signed registry with 2-of-3 trust-root rotation is built; what remains is its first publication — an irreversible key ceremony, written up as a runbook rather than a document. It is the only ACP work still open.
Once the repository is public, every binary carries a verifiable attestation:
gh attestation verify flowcodex-linux-x64 --repo Msvnc0/flowcodex.
The deliberately deferred list, kept from the decision records: cloud sync and memory portability, session sharing, a GitHub App, and a VS Code extension. All of them are unblocked by the headless server; none of them will ship before the core gates stay green.
A GitHub Action opens a models.dev catalog refresh PR every Monday. The truth tests keep docs and settings honest against the source. Version 1.0 is not a feature date — it arrives when the architecture, contract and coverage gates hold steady.