A Block-graph storage engine for LLM Agents
Apache-2.0 · R6RS Scheme, Chez as the reference platform · runs on Igropyr
To get past these limits, the system stops organizing content by files. Blocks now store both documentation and actual source code files. The underlying store still projects onto standard files when needed, but a file boundary no longer decides what is read, what is written, or what is run. Instead, we rebuilt how engineering knowledge and code are structured, assembled, tested, and computed across seven core mechanisms.
We built a compact interface strictly for machine economics and token budgets, stripping out rendering markup entirely. Core verbs like outline, search, refs, and batch answer in dense S-expressions, one item per line. An agent can see the exact token footprint of what it is about to read before reading it—and there is no second, pretty-printed output mode to parse.
$ theourgia outline --depth 1
- 8pniyrtk.j Home
- 8pniyrtk.k Why not markdown
- 8pniyrtk.l Model
$ theourgia search "stale-baseline"
(hit "8pniyrtk.2l" 4 "refusal, stale-baseline, commit, example, answer")
(hit "8pniyrtk.1b" 1 "A refused commit does not simply say no. It answers stale-baseline, and it names")The wall between the code repository and the document store is a filesystem accident, not a property of the work itself. Here, blocks hold both Markdown documentation and source code files—everything engineering produces lives in one unified graph. A design section, a full source file, a Scheme macro, a Python function, and an architectural decision record are all equal blocks, each with its own stable ID and each independently addressable. Edges are typed and directed: a requirement points to the code that meets it, the code points to the test that checks it, and the test points back to the discussion that decided it.
In a Scheme library block this is literal rather than a metaphor. A definition is a block, and evaluating an expression against that library is evaluating the definitions those blocks hold -- so a name resolves to its definition block, log <id> gives that definition's history, and an explicit edge ties it to the design that motivated it. The same three questions about one name -- what it is, how it got that way, why it exists -- are three reads rather than three searches.
Source in other languages goes in as text blocks cut at top-level regions, with names extracted as well as the language table allows. They are blocks like any other: addressable, linkable, with their own history. What they do not get is the resolution above -- a name in Python source is not a binding the store can follow to a block.
The graph changes how an agent—or a human—hunts for context, and what that hunt costs. A single refs query answers the whole picture at once: which design mentions this function? which prose block refers to it? Meanwhile, log hands back the block's exact history, and search pinpoints the answering block so the agent reads only that block instead of the entire file around it.
$ theourgia link koou2buo.2 implements koou2buo.1
$ theourgia insert --under root --title "Mentions by id" \
--text "See [[koou2buo.1]] for the reasoning."
$ theourgia refs koou2buo.1
(ok (items (ref (from "koou2buo.2") (rel implements) (via link))
(ref (from "koou2buo.4") (rel ref) (via md))))Measured across 56 questions over 20 real engineering documents: searching and reading the target block costs 15% of the tokens required to read the full document each time; browsing the outline before reading costs 20%. The outline of an entire corpus takes up just 6% of its raw Markdown size.
| Route | Tokens, 56 questions | Share |
|---|---|---|
| Read the whole document for each question | 232,478 | 100% |
| outline, then read the block that answers | 46,410 | 20% |
| search, then read the block that answers | 34,917 | 15% |
An agent's real bottleneck isn't I/O speed—it's whether the tens of thousands of bytes entering its context window are actually right and complete. Today, agents assemble context by intuition: run a search, read a few blocks, feel like it's enough, and start coding. If they miss a premise or follow a superseded ruling, the mistake only surfaces after the work is done. We turn "gathering materials" from an agent's heuristic into a deterministic store verb—and make the store accountable for its choices.
It works in two layers:
The store isn't just for the automated half; it is a shared semantic layer where humans and agents work on the exact same code blocks, prose blocks, edges, and history. For humans, that means an IDE, not a CLI. Because people edit real files projected by the store, the Language Server works natively with zero loss of IntelliSense—while the extension maps every save back to the blocks it touched and adds what only the graph knows:
Multiple agents and humans can edit and test the exact same code or documentation block simultaneously without locking up the store or stepping on each other's runtime state.
$ theourgia commit <block> --writer w1
(error stale-baseline
(block "s9av8p0w.1")
(based-on "fb580a74...")
(now "904e0323...")
(since (("s9av8p0w" . 3)
(<the record that landed>)
(set "s9av8p0w.1" src "w2 version")))
(usage (commit (<block> ...) ("--writer" <name>)
("--working-version" <block>=<version>))))The store is not just where knowledge and code are kept; it is a runtime they execute in.
$ theourgia eval --under <library> '(f)'
(ok (values ("second version")) (working-view #f (("koou2buo" . 8))) ...)
$ theourgia eval --under <library> '(f)' --cut '(("koou2buo" . 7))'
(ok (values ("committed")) (working-view #f (("koou2buo" . 7))) ...)The definition was changed between the two calls. The second asks for the earlier cut and gets the behaviour that was there then.
# a store is a directory
theourgia init
# one block answers one question
theourgia insert --under root --title "Why the retry is three deep" \
--keywords "retry, supervisor, worker, timeout" \
--text "..."
# find it again without reading the store
theourgia search "retry supervisor"
theourgia read <id>
# what is under this block?
theourgia outline --depth 1Theourgia is R6RS Scheme and builds with Chez Scheme against Igropyr. build.ss compiles every library in the tree into a directory of objects; the client, the MCP shell and the sandbox worker are the scripts beside them. There is no package to install and no daemon to configure: point the client at a directory and it starts what it needs.