T

Theourgia

A Block-graph storage engine for LLM Agents

Apache-2.0 · R6RS Scheme, Chez as the reference platform · runs on Igropyr

Core features

A Living Graph Engine

To get past these limits, the system stops organizing content by files. Blocks now store both documentation and actual source code files. The underlying store still projects onto standard files when needed, but a file boundary no longer decides what is read, what is written, or what is run. Instead, we rebuilt how engineering knowledge and code are structured, assembled, tested, and computed across seven pieces.

01

A Minimal Query Surface for Agents

We built a compact interface for token budgets, and stripped out the rendering markup. Core verbs like outline, search, refs, and batch answer in dense S-expressions, one item per line. An agent can see the token footprint of what it is about to read before reading it—and there is no second, pretty-printed output mode to parse.

$ theourgia outline --depth 1
- 8pniyrtk.j  Home
- 8pniyrtk.k  Why not markdown
- 8pniyrtk.l  Model

$ theourgia search "stale-baseline"
(hit "8pniyrtk.2l" 4 "refusal, stale-baseline, commit, example, answer")
(hit "8pniyrtk.1b" 1 "A refused commit does not simply say no. It answers stale-baseline, and it names")
02

One Unified Graph Across Docs and Source Code

The wall between the code repository and the document store is a filesystem accident, not a property of the work itself. Here, blocks hold both Markdown documentation and source code files—everything engineering produces lives in one unified graph. A design section, a full source file, a Scheme macro, a Python function, and an architectural decision record are all equal blocks, each with its own stable ID and each independently addressable. Edges are typed and directed, so you can walk from a requirement to the code that meets it, to the test that checks it, back to the discussion that decided it.

refrootthe storeModelf0sjkuzn.3Concurrencyf0sjkuzn.7Agentsf0sjkuzn.bA blockf0sjkuzn.4Drafts, then commitf0sjkuzn.8The write protocolf0sjkuzn.c

The store, the block and the log, in full ->

03

Global Backlinks at Near-Zero Cost

The graph changes how an agent—or a human—hunts for context, and what that hunt costs. A single refs query answers the whole picture at once: which design mentions this function? which prose block refers to it? log gives you the block's full history, and search lands on the one block that answers—so the agent reads that block, not the whole file around it.

$ theourgia link koou2buo.2 implements koou2buo.1
$ theourgia insert --under root --title "Mentions by id" \
    --text "See [[koou2buo.1]] for the reasoning."

$ theourgia refs koou2buo.1
(ok (items (ref (from "koou2buo.2") (rel implements) (via link))
           (ref (from "koou2buo.4") (rel ref) (via md))))

Measured across 56 questions over 20 real engineering documents: searching and reading the target block costs 15% of the tokens required to read the full document each time; browsing the outline before reading costs 20%. The outline of an entire corpus takes up just 6% of its raw Markdown size.

RouteTokens, 56 questionsShare
Read the whole document for each question232,478100%
outline, then read the block that answers46,41020%
search, then read the block that answers34,91715%
04

Deterministic Context Assembly

An agent's real bottleneck isn't I/O speed—it's whether the tens of thousands of bytes entering its context window are actually right and complete. Today, agents assemble context by intuition: run a search, read a few blocks, feel like it's enough, and start coding. If they miss a premise or follow a superseded ruling, the mistake only surfaces after the work is done. We turn "gathering materials" from an agent's heuristic into a deterministic store verb—and make the store show its work.

It works in two layers:

05

Humans and Agents in the IDE

The store isn't just for the automated half; it is a shared semantic layer where humans and agents work on the same code blocks, prose blocks, edges, and history. For humans, that means an IDE, not a CLI. Because people edit real files projected by the store, the Language Server works natively without losing IntelliSense—while the extension maps every save back to the blocks it touched and adds what only the graph knows:

06

Zero-Conflict Concurrent Workspaces & Isolated Execution

Multiple agents and humans can edit and test the exact same code or documentation block simultaneously without locking up the store or stepping on each other's runtime state.

$ theourgia commit <block> --writer w1

(error stale-baseline
  (block "s9av8p0w.1")
  (based-on "fb580a74...")
  (now "904e0323...")
  (since (("s9av8p0w" . 3)
          (<the record that landed>)
          (set "s9av8p0w.1" src "w2 version")))
  (usage (commit (<block> ...) ("--writer" <name>)
                 ("--working-version" <block>=<version>))))
07

A Lisp Machine: A Living Knowledge Base

The store keeps knowledge and code — and runs them, too.

$ theourgia eval --under <library> '(f)'
(ok (values ("second version")) (working-view #f (("koou2buo" . 8))) ...)

$ theourgia eval --under <library> '(f)' --cut '(("koou2buo" . 7))'
(ok (values ("committed")) (working-view #f (("koou2buo" . 7))) ...)

The definition was changed between the two calls. The second asks for the earlier cut and gets the behaviour that was there then.

Measured

Numbers, and where they came from

387blocks written into one store by five agents at the same time, with 0 refusals
5.7xless context for one real question: outline plus one block, against reading the file
113/113documents exported back byte-for-byte identical after being imported
6%of a corpus's markdown tokens is the size of its entire outline
1 msto read a block through the daemon
How the numbers were taken Every figure on this site comes from one of two measured runs: 157 files of a real agent memory written by five concurrent agents, and 113 documents converted and then queried. Four things are worth knowing before quoting them. The token counts use cl100k_base, which is OpenAI's tokenizer, not Claude's and not a billing unit; calibrated once against real context growth it ran about 9% low. How much a query saves tracks the size of the block that happens to answer it, not anything theourgia does: in the bucket where the answering block is most of the document the median query still cost 95.2% of reading the whole thing. Small documents have little to give either -- under about a thousand tokens the median was 38.5%, and one README's outline alone came to 67.6% of the document it summarised. And roughly 40% of queries made without knowing what is in the store fall back to reading the outline first.
Start

A store is a directory

# a store is a directory
theourgia init

# one block answers one question
theourgia insert --under root --title "Why the retry is three deep" \
  --keywords "retry, supervisor, worker, timeout" \
  --text "..."

# find it again without reading the store
theourgia search "retry supervisor"
theourgia read <id>

# what is under this block?
theourgia outline --depth 1

Theourgia is R6RS Scheme and builds with Chez Scheme against Igropyr. build.ss compiles every library in the tree into a directory of objects; the client, the MCP shell and the sandbox worker are the scripts beside them. There is no package to install and no daemon to configure: point the client at a directory and it starts what it needs.