θ

Theourgia

A Block-graph storage engine for LLM Agents

Apache-2.0 · R6RS Scheme, Chez as the reference platform · runs on Igropyr

Core features

A Living Graph Engine for Human-Agent Symbiosis

To get past these limits, the system stops organizing content by files. Blocks now store both documentation and actual source code files. The underlying store still projects onto standard files when needed, but a file boundary no longer decides what is read, what is written, or what is run. Instead, we rebuilt how engineering knowledge and code are structured, assembled, tested, and computed across seven core mechanisms.

01

A Minimal Query Surface for Agents

We built a compact interface strictly for machine economics and token budgets, stripping out rendering markup entirely. Core verbs like outline, search, refs, and batch answer in dense S-expressions, one item per line. An agent can see the exact token footprint of what it is about to read before reading it—and there is no second, pretty-printed output mode to parse.

$ theourgia outline --depth 1
- 8pniyrtk.j  Home
- 8pniyrtk.k  Why not markdown
- 8pniyrtk.l  Model

$ theourgia search "stale-baseline"
(hit "8pniyrtk.2l" 4 "refusal, stale-baseline, commit, example, answer")
(hit "8pniyrtk.1b" 1 "A refused commit does not simply say no. It answers stale-baseline, and it names")
02

One Unified Graph Across Docs and Source Code

The wall between the code repository and the document store is a filesystem accident, not a property of the work itself. Here, blocks hold both Markdown documentation and source code files—everything engineering produces lives in one unified graph. A design section, a full source file, a Scheme macro, a Python function, and an architectural decision record are all equal blocks, each with its own stable ID and each independently addressable. Edges are typed and directed: a requirement points to the code that meets it, the code points to the test that checks it, and the test points back to the discussion that decided it.

refrootthe storeModelf0sjkuzn.3Concurrencyf0sjkuzn.7Agentsf0sjkuzn.bA blockf0sjkuzn.4Drafts, then commitf0sjkuzn.8The write protocolf0sjkuzn.c

A name is a block

In a Scheme library block this is literal rather than a metaphor. A definition is a block, and evaluating an expression against that library is evaluating the definitions those blocks hold -- so a name resolves to its definition block, log <id> gives that definition's history, and an explicit edge ties it to the design that motivated it. The same three questions about one name -- what it is, how it got that way, why it exists -- are three reads rather than three searches.

Other languages, more modestly

Source in other languages goes in as text blocks cut at top-level regions, with names extracted as well as the language table allows. They are blocks like any other: addressable, linkable, with their own history. What they do not get is the resolution above -- a name in Python source is not a binding the store can follow to a block.

The store, the block and the log, in full ->

03

Global Backlinks at Near-Zero Cost

The graph changes how an agent—or a human—hunts for context, and what that hunt costs. A single refs query answers the whole picture at once: which design mentions this function? which prose block refers to it? Meanwhile, log hands back the block's exact history, and search pinpoints the answering block so the agent reads only that block instead of the entire file around it.

$ theourgia link koou2buo.2 implements koou2buo.1
$ theourgia insert --under root --title "Mentions by id" \
    --text "See [[koou2buo.1]] for the reasoning."

$ theourgia refs koou2buo.1
(ok (items (ref (from "koou2buo.2") (rel implements) (via link))
           (ref (from "koou2buo.4") (rel ref) (via md))))

Measured across 56 questions over 20 real engineering documents: searching and reading the target block costs 15% of the tokens required to read the full document each time; browsing the outline before reading costs 20%. The outline of an entire corpus takes up just 6% of its raw Markdown size.

RouteTokens, 56 questionsShare
Read the whole document for each question232,478100%
outline, then read the block that answers46,41020%
search, then read the block that answers34,91715%
What refs resolves, exactly Measured, because this is easy to over-claim. A mention resolves by id: a block whose text contains [[koou2buo.1]] shows up as (ref (from ...) (rel ref) (via md)). A mention by title does not -- a block written with [[The design]] did not appear in refs for the block titled The design. An explicit edge shows up as (ref (from ...) (rel implements) (via link)) under whatever relation name it was given. So renaming a block leaves explicit edges and id mentions intact, because both are by id; what does not survive a rename is a reference that was only ever a name.
04

Deterministic Context Assembly

An agent's real bottleneck isn't I/O speed—it's whether the tens of thousands of bytes entering its context window are actually right and complete. Today, agents assemble context by intuition: run a search, read a few blocks, feel like it's enough, and start coding. If they miss a premise or follow a superseded ruling, the mistake only surfaces after the work is done. We turn "gathering materials" from an agent's heuristic into a deterministic store verb—and make the store accountable for its choices.

It works in two layers:

05

Human-Agent Symbiosis in the IDE

The store isn't just for the automated half; it is a shared semantic layer where humans and agents work on the exact same code blocks, prose blocks, edges, and history. For humans, that means an IDE, not a CLI. Because people edit real files projected by the store, the Language Server works natively with zero loss of IntelliSense—while the extension maps every save back to the blocks it touched and adds what only the graph knows:

06

Zero-Conflict Concurrent Workspaces & Isolated Execution

Multiple agents and humans can edit and test the exact same code or documentation block simultaneously without locking up the store or stepping on each other's runtime state.

$ theourgia commit <block> --writer w1

(error stale-baseline
  (block "s9av8p0w.1")
  (based-on "fb580a74...")
  (now "904e0323...")
  (since (("s9av8p0w" . 3)
          (<the record that landed>)
          (set "s9av8p0w.1" src "w2 version")))
  (usage (commit (<block> ...) ("--writer" <name>)
                 ("--working-version" <block>=<version>))))
About that example That answer was taken from a run, not written by hand: two writers took a draft of one block from the same baseline, the second committed, and the first was refused. The two hashes are shortened here to fit the page, and the record in the middle of since is stood in for by a description of it; everything else is as the store answered.
07

A Lisp Machine: A Living Knowledge Base

The store is not just where knowledge and code are kept; it is a runtime they execute in.

$ theourgia eval --under <library> '(f)'
(ok (values ("second version")) (working-view #f (("koou2buo" . 8))) ...)

$ theourgia eval --under <library> '(f)' --cut '(("koou2buo" . 7))'
(ok (values ("committed")) (working-view #f (("koou2buo" . 7))) ...)

The definition was changed between the two calls. The second asks for the earlier cut and gets the behaviour that was there then.

Measured

Numbers, and where they came from

387blocks written into one store by five agents at the same time, with 0 refusals
5.7xless context for one real question: outline plus one block, against reading the file
113/113documents exported back byte-for-byte identical after being imported
6%of a corpus's markdown tokens is the size of its entire outline
1 msto read a block through the daemon
How the numbers were taken Every figure on this site comes from one of two measured runs: 157 files of a real agent memory written by five concurrent agents, and 113 documents converted and then queried. Four things are worth knowing before quoting them. The token counts use cl100k_base, which is OpenAI's tokenizer, not Claude's and not a billing unit; calibrated once against real context growth it ran about 9% low. How much a query saves tracks the size of the block that happens to answer it, not anything theourgia does: in the bucket where the answering block is most of the document the median query still cost 95.2% of reading the whole thing. Small documents have little to give either -- under about a thousand tokens the median was 38.5%, and one README's outline alone came to 67.6% of the document it summarised. And roughly 40% of queries made without knowing what is in the store fall back to reading the outline first.
Start

A store is a directory

# a store is a directory
theourgia init

# one block answers one question
theourgia insert --under root --title "Why the retry is three deep" \
  --keywords "retry, supervisor, worker, timeout" \
  --text "..."

# find it again without reading the store
theourgia search "retry supervisor"
theourgia read <id>

# what is under this block?
theourgia outline --depth 1

Theourgia is R6RS Scheme and builds with Chez Scheme against Igropyr. build.ss compiles every library in the tree into a directory of objects; the client, the MCP shell and the sandbox worker are the scripts beside them. There is no package to install and no daemon to configure: point the client at a directory and it starts what it needs.