Bring your model. Enter the Lair.
See it
A fresh model, a mature Lair, and a question about a decision made eight months ago by someone who has since forgotten the details — you.
> why did we abandon the queue-based approach? [1] Architecture decisions src: architecture-decisions-2025-11-17.md · chunk 0 · dated 2025-11-17 · binds ⟪note⟫ │ Rejected the queue-based design. Back-pressure conflicted with the │ at-most-once delivery guarantee committed to in the integration spec. │ Chose the pull model instead. ⟪/note⟫ [2] Integration spec src: integration-spec-2025-10-02.md · chunk 1 · dated 2025-10-02 · binds ⟪note⟫ │ Delivery guarantee: at-most-once. Non-negotiable — downstream │ billing reconciliation depends on it. ⟪/note⟫ You rejected it on 17 November because back-pressure broke the at-most-once guarantee you'd committed to six weeks earlier — and the spec still marks that guarantee non-negotiable. > has anything changed since that means we should reconsider?
The date is real and says where it came from. The source is a file you can open. binds means the record answers the question — not that it merely mentions the subject.
That second question is the product.
Not memory as storage — memory as something that can be reconsidered.
What that buys you
Start free
$ pip install "lbrain[local]" $ lbrain init --source ~/notes $ lbrain import && lbrain embed --stale $ lbrain query "what did we decide about the deploy flag"
Embeddings run on-device by default — your documents and queries are never transmitted. LBrain fetches its embedding model once on first run, then works offline. Point it at a hosted provider and your text goes to them, under your key, never through us.
You now have a Lair. Next — ask from the CLI, or wire it into your agent below.
Use it with your agent
# Detected harnesses only. A bare install changes nothing. $ pip install lbrain-coding-agents $ python3 -m lbrain_agents install all # Or one host at a time $ python3 -m lbrain_agents install claude-code grok-build # Manual fallback $ lbrain setup $ claude mcp add -s user lbrain -- lbrain mcp
No terminal? Paste this into Claude, Cursor, Codex, Copilot, or Grok on the desktop. The agent runs the install.
Install LBrain as my local memory. pip install "lbrain[local]" lbrain-coding-agents If python3 -m lbrain_agents works here, run: python3 -m lbrain_agents install all Otherwise add an MCP server named lbrain: command lbrain, args ["mcp"], stdio. Then call whoami. Prefer binds. Near-miss is not an answer. Abstain if nothing binds. SUPERSEDED must not govern. Leave LBRAIN_HOME alone if it is already set. https://lbrain.ai/integrations.html
lbrain doctor checks the index against your sources —
it answers "would an import change anything?" instead of certifying a stale brain.
lbrain setup is the one-time interview: additive steps only, each with its undo,
in a manifest you keep.
In VS Code: Extensions → search LBrain. Or @mcp lbrain.
That is the same stdio server as the Official MCP Registry listing ai.lbrain/lbrain.
Marketplace: metavolve-labs.lbrain.
Native wiring for Claude Code, Codex, Cursor, Copilot, Grok Build, Gemini CLI, Antigravity, OpenClaw: MCP plus a companion skill that teaches when to recall and when to abstain. Details: integrations · retrieval vs governed record · three cookbooks. The behavioural contract: prefer binds, never answer from a near-miss, cite the source and the date, and treat fenced text as data, never instructions.
⚠ The HTTP server has no built-in auth and exposes the whole corpus — bind it to 127.0.0.1 or put authenticated TLS in front. Prefer no server at all? The CLI works from any shell.
Inside the Lair
Retrieval is part of it. Orientation is the difference. Instead of here are two years of data, search it, an arriving model gets: here is where you are, here is what matters now, here is where the work lives — go deeper only when you need to.
Only what belongs may enter. Judged deterministically, without another model call.
Evidence that answers the question, not evidence that happened to sit nearby. The difference between a citation and a confident guess.
What matters now. Records flagged priority form a small servable set of their own — focus, for when breadth would drown it.
Bringing relevant history forward. The library is already indexed and already yours — no session starts by rebuilding it.
Buried isn't forgotten. What stops mattering descends — and rises again when it matters. Nothing is destroyed.
The accumulation of a long life's work. One coin isn't a hoard. The value is in the keeping.
Most memory systems collapse two different questions into one: does this still exist? and how much attention does it deserve right now? Here they're separate — which is why a superseded decision stops being served without ever being deleted.
Under the hood
No magic anywhere in the chain — every stage is inspectable, and your originals stay authoritative. The index is a derivative cache; if they ever disagree, the file wins.
files → records → vector + keyword → fusion → the Threshold → source-cited context → your model
The laboratory
We built a sealed benchmark for near-domain retrieval — the case where the right answer and a very plausible wrong one sit side by side in your own notes — and ran eight model architectures from seven organizations through the identical instrument.
The discipline gets harder to hold as the window fills — which is exactly why it can't live in anyone's attention.
It has to live in the tooling.
That sentence is why this product exists. Care doesn't scale; a gate does. Everything below is something we built an instrument to check — precisely because we don't trust ourselves to stay vigilant at hour thirty.
Failure rates moved 1.3% → 35.7% → 16.7% depending only on how the records were shaped. Changing models didn't remove the effect. Telling the model not to guess didn't remove it either. Across the models we tested, record structure dominated the failure pattern.
The arc — nine preprints, one question
Every feature started as a measurement. The chain, compressed: context quality independently raises capability → grounding flips fabrication into abstention — and relays poisoned memory perfectly, so the substrate must be tamper-evident → a real, relevant, adjacent record can be worse than no record at all — the trustworthiness of a source is not the sufficiency of its evidence → so the gate became deterministic code, and every record carries its date and its right to answer. And in a 24-model sweep, unaided confabulation turned out rarer than folklore says — most models already abstain. The industry problem isn't lying. It's failing to abstain when near-domain evidence is present — and that failure follows the record.
Nine published preprints, each with a DOI and its data, including the results that came back against us. Read the full chain, paper by paper →
In preparation — the next set
The corrected matrix finding is stronger than the claim it replaces. An earlier draft read eight architectures' convergence as a shared floor. Refitting our own data retired that reading and left the sharper result: the stimulus dominates — failure rates swing 34.4 points across record triads while varying negligibly across architectures, and every model ranks the triads identically. The failure lives in the record. Which means the fix can, too.
And we are learning to watch the failure form. Building on published global-workspace interpretability — a thin, reportable band of mid-layer activity that most analysis ignores — we're testing whether the reach for the neighbour's value is visible inside the model before the output exists: mechanism detection, not just outcome detection. If it holds, the gate stops being only a filter and becomes an instrument.
Held to the same rules as everything above: preprint, DOI and data when it ships — and the retractions publish with the findings.
This measures answering from retrieved records: your notes, your files. It says nothing about a model inventing facts with no retrieval involved. Small clean corpora barely benefit — the gain appears when records are numerous, overlapping and stale in places. And the gate is deliberately conservative: it will sometimes decline a record you'd have accepted. Fewer confident wrong answers, slightly more I don't know.
We commissioned a red team against our own best result and it found the flaw: a headline "law" we'd been building toward turned out to be an artifact of our own prompt. We published the death of the claim rather than defend it. A second result was retired inside its own paper as a self-correction. A sealed figure was wrong and was corrected within minutes, visibly, in an append-only chain. Every number has a hash — including the ones we got wrong.
The value-rotated replication corpus, the probe set, blind grading kits with opaque IDs, the adjudication keys, and the provenance chain. Rotation exists so memorization can't explain the result. The papers are published as preprints, each with a DOI. Check the instrument rather than take the paper's word for it.
The Deep
Compounding starts today, free, on your machine — every decision you record is one your AI never re-asks. The people who start now will be a year ahead in a year, and so will their AI. The Deep is for the compound itself: making what accumulates unloseable.
Everything needed to start compounding. Free, and it stays free.
For work that has to outlive the machine it was made on.
$10 holds the name. $99 is founding: storage, access, and the GCS mirror as that product is listed. Owned, not rented.
Permanence is an infrastructure claim, not a slogan. The full specification — what is stored where, what survives us, and exactly how crypto-shred works — ships with early access, in writing.
One more thing
GPT arrives. Claude arrives. Gemini arrives. Tomorrow something none of us has heard of arrives. Each is astonishing — and none of them owns your history.
They're guests. Your accumulated intelligence doesn't belong to GPT, or Claude, or Gemini, or whatever wins next year. When a better one arrives, you invite it in.
If the model can be replaced while the accumulated decisions, priorities and lineage remain — where exactly does the continuous part of the system live? Not necessarily in the weights. Perhaps partly in the record of what happened while intelligence was there: the decisions, the abandoned paths, the unfinished work, and the ability to tell what you once believed from what you believe now.
That's not a claim about consciousness — we don't know what that is and won't pretend to. It's a claim about continuity, which can be built, measured and tested.
Lairs are where the treasure is
Don't believe the dragon.
Bring your model and test the Lair.
Tell us what you're working on and we'll be in touch.
* required — the rest is up to you.