Markdown and plain source files are great for humans to write in, but as the underlying memory and engineering substrate for multi-agent systems, physical files hit five systemic walls.
Plain text offers zero structural validation on write. When an agent makes a bad edit or a botched patch, it fails silently — quietly dropping a design section or mangling code without throwing an error. And if the agent reads the file back after every write just to verify its own work, it doubles the token burn.
Traditional files have no block-level isolation. When multiple agents and humans reason in parallel and edit the same docs or source files, they blindly overwrite each other's work — and if they try to build or test in a shared working directory, they corrupt each other's execution state.
Beyond the cost of write-verification, everyday reading is equally wasteful. Docs and source files tend to grow monotonically — N lines added, zero deleted. Because the file is the unit of addressing, an agent that only needs one specific rule or function is forced to read the entire monolith, burning through its context window.
Measured on one real question against one real memory file, both routes run into the same agent's context: reading the 19,820-byte file cost 10,432 tokens; an outline call plus a read of the single block that answered the question cost 1,823. That is 5.7x, and about 600 of each route's tokens are per-turn framing rather than content.
| Route | Tokens | Share |
|---|---|---|
| Read the whole document for each question | 232,478 | 100% |
| outline, then read the block that answers | 46,410 | 20.0% |
| search, then read the block that answers | 34,917 | 15.0% |
Fifty-six questions over twenty documents. The share is of what reading the whole document every time would have cost.
Over time, a file-based knowledge base inevitably decays into two extremes: thousands of isolated, atomic notes, and a handful of massive files accumulated over months. Splitting those monoliths by hand only shatters the logical flow.
External vector indexes and line-number mappings are detached from the text itself. The moment a paragraph or function moves, the index goes stale without warning — feeding hallucinations to the agent while leaving cross-file semantic breaks invisible to Git.
Worth being precise, because it is easy to over-claim. Explicit edges are recorded by id, so renaming a block's title or a function's name leaves them intact. Derived edges -- the ref and uses relations the reducer works out for itself -- are re-resolved by name on every reduction. That is a fresh resolution, not an edge that followed the rename, and after a rename it resolves to whatever now carries the name.
Changing one sentence in a markdown memory means handing the whole file back. Here it means handing back the one block: write <block> <bytes>, then commit. Measured on three documents, the payload plus command line came to between 0.19% and 5.61% of rewriting the file, and the answer stayed at 190-200 tokens whatever the document's size. Exporting afterwards and diffing against the original showed exactly one line replaced, all three times.
| Document | Rewrite the file | write plus commit | Share |
|---|---|---|---|
| 19,820 B design note | 8,403 | 292 | 3.47% |
| 367,298 B design note | 127,510 | 246 | 0.19% |
| 10,624 B design note | 2,904 | 163 | 5.61% |