Sage Nord

← All posts

Cortyx: why your coding agent shouldn't re-read the whole repo on every turn

A long-lived AI coding session has a context problem that gets worse the longer it runs. Either the agent re-reads large chunks of the codebase on every turn, which burns tokens and blows past prompt-cache boundaries, or a memory layer decides what to inject instead, which usually means calling another LLM to summarize or rank candidate memories before the actual task can start. Both approaches pay a real cost on every single turn.

Cortyx, an open-source Rust project from Sage Nord, is built to avoid that trade-off entirely. It's an MCP-native context delivery engine: instead of synthesizing an answer about what's relevant, it retrieves pre-written, human-readable context and hands it to the agent, which does the synthesis itself. No LLM call in the retrieval path, no cloud backend, and a retrieval-only activation budget of about 40ms.

The problem: context delivery, disguised as memory

Most "AI memory" systems conflate two different jobs: deciding what's relevant, and explaining what's relevant. Doing both well usually means running the candidate context back through an LLM, which is accurate but slow (often close to a second, dominated by inference) and requires a hosted model. That's a reasonable trade for some products. It's a bad one for a coding agent that needs to make this decision dozens of times in a single session, and it's a non-starter for anyone who wants their project's context to stay fully local.

Cortyx's position: separate the two jobs. Retrieval should be fast, deterministic, and local. Synthesis is what the LLM you're already paying for is good at, so let it do that part.

How it works

Cortyx runs a two-stage pipeline: an indexing stage that runs once per code change, and a query stage that runs on every agent task.

Cortyx two-stage pipeline: index time and query time Index time: project source is AST-parsed into human-readable neuron files, then built into a hybrid BM25 plus dense-embedding index. Query time: an agent task triggers hybrid retrieval and reranking, followed by synapse graph traversal up to three hops, producing 3-5 ranked neurons that are injected after the prompt-cache breakpoint into the LLM's context window, keeping the cached prefix byte-identical. INDEX TIME (once per change) Project source AST-parsed by cortyx compile Neurons .cortyx/neurons/*.md — plain, git-tracked Hybrid index BM25 + dense embeddings QUERY TIME (≤40ms activation, per task) Agent task "add dark mode to SwiftUI view" Retrieval + rerank hybrid BM25 + dense, top-k Synapse traversal graph walk, up to 3 hops 3–5 neurons selected, ordered by relevance injected after the prompt-cache breakpoint — static prefix stays byte-identical LLM context window (cache-hit)
Cortyx's two-stage pipeline: neurons are built once at index time, then retrieved and ranked in under 40ms at query time, landing after the prompt-cache breakpoint.

At index time, cortyx compile walks the project and AST-parses it into neurons: plain Markdown files under .cortyx/neurons/, one per source file or concept, git-tracked and human-editable like any other file in the repo. There's no proprietary database to inspect or lose.

At query time, a task like "add dark mode to a SwiftUI view" triggers hybrid retrieval (BM25 keyword search plus dense embeddings, reranked), followed by synapse graph traversal up to three hops to pull in related neurons a pure similarity search would miss. The result is 3–5 ranked neurons, injected after the prompt-cache breakpoint, so the static prefix of the prompt stays byte-identical across turns and provider caches keep hitting instead of resetting.

That last detail is the one that's easy to miss and actually matters most for cost: a context system that changes the beginning of the prompt on every turn defeats prompt caching everywhere it's used, quietly turning every turn into an uncached (and much more expensive) call.

The results

On the LME-500 long-memory benchmark, Cortyx reports 96.8% macro R@5 retrieval recall, ahead of MemPalace's 96.6% on the same retrieval-recall metric:

Retrieval recall (R@5) on the LME-500 benchmark Bar chart of macro R@5 retrieval recall on LME-500: Cortyx 96.8%, MemPalace 96.6%, mem0 v3 94.6% (mem0's figure is LLM-as-judge answer accuracy, a related but distinct metric, shown here for scale). 90% 92% 94% 96% 98% 100% Cortyx 96.8% MemPalace 96.6% mem0 v3 94.6%
Macro R@5 retrieval recall on LME-500. Cortyx and MemPalace report the same retrieval-recall metric; mem0's figure is LLM-as-judge answer accuracy, a related but different measurement, shown for scale.

The comparison to mem0 needs its own caveat, which the project documents directly rather than glossing over: mem0 reports LLM-as-judge answer accuracy (94.4–94.8%), an end-to-end metric that includes an LLM call, while Cortyx and MemPalace report retrieval recall, a narrower, retrieval-only metric. They're related but not the same measurement. Where the comparison is closer to apples-to-apples is latency and infrastructure: Cortyx's retrieval-only activation is roughly 22ms at p95, against mem0's full pipeline at roughly 900ms–1.1s p50 because it includes an LLM call, and Cortyx requires no runtime model at all, running fully local and offline. The project ships a cortyx proof-certificate command specifically so these claims can be independently reproduced rather than taken on faith.

Where it fits

Cortyx installs with a one-line script or cargo install cortyx, indexes a project with cortyx compile, and starts an MCP server that Claude Code, Cursor, Codex, Windsurf, VS Code, and Zed can all connect to out of the box, with cortyx install auto-configuring whichever of those it detects. Beyond retrieval, it layers on a knowledge graph for temporal facts, per-agent diaries for session handoff, and a "concepts" library for sharing proven patterns across projects, all stored as the same local, git-tracked Markdown.

The trade-off is the same one that produces its speed: Cortyx delivers context, it doesn't generate answers about your code. Teams that want a system to summarize or reason over retrieved context on their behalf still need an LLM call somewhere in the loop, Cortyx just makes sure that call isn't spent re-discovering what the agent already knew last session.

Source: sorunokoe/Cortyx on GitHub.