Cortyx: why your coding agent shouldn't re-read the whole repo on every turn
A long-lived AI coding session has a context problem that gets worse the longer it runs. Either the agent re-reads large chunks of the codebase on every turn, which burns tokens and blows past prompt-cache boundaries, or a memory layer decides what to inject instead, which usually means calling another LLM to summarize or rank candidate memories before the actual task can start. Both approaches pay a real cost on every single turn.
Cortyx, an open-source Rust project from Sage Nord, is built to avoid that trade-off entirely. It's an MCP-native context delivery engine: instead of synthesizing an answer about what's relevant, it retrieves pre-written, human-readable context and hands it to the agent, which does the synthesis itself. No LLM call in the retrieval path, no cloud backend, and a retrieval-only activation budget of about 40ms.
The problem: context delivery, disguised as memory
Most "AI memory" systems conflate two different jobs: deciding what's relevant, and explaining what's relevant. Doing both well usually means running the candidate context back through an LLM, which is accurate but slow (often close to a second, dominated by inference) and requires a hosted model. That's a reasonable trade for some products. It's a bad one for a coding agent that needs to make this decision dozens of times in a single session, and it's a non-starter for anyone who wants their project's context to stay fully local.
Cortyx's position: separate the two jobs. Retrieval should be fast, deterministic, and local. Synthesis is what the LLM you're already paying for is good at, so let it do that part.
How it works
Cortyx runs a two-stage pipeline: an indexing stage that runs once per code change, and a query stage that runs on every agent task.
At index time, cortyx compile walks the project and AST-parses it into neurons: plain Markdown files under .cortyx/neurons/, one per source file or concept, git-tracked and human-editable like any other file in the repo. There's no proprietary database to inspect or lose.
At query time, a task like "add dark mode to a SwiftUI view" triggers hybrid retrieval (BM25 keyword search plus dense embeddings, reranked), followed by synapse graph traversal up to three hops to pull in related neurons a pure similarity search would miss. The result is 3–5 ranked neurons, injected after the prompt-cache breakpoint, so the static prefix of the prompt stays byte-identical across turns and provider caches keep hitting instead of resetting.
That last detail is the one that's easy to miss and actually matters most for cost: a context system that changes the beginning of the prompt on every turn defeats prompt caching everywhere it's used, quietly turning every turn into an uncached (and much more expensive) call.
The results
On the LME-500 long-memory benchmark, Cortyx reports 96.8% macro R@5 retrieval recall, ahead of MemPalace's 96.6% on the same retrieval-recall metric:
The comparison to mem0 needs its own caveat, which the project documents directly rather than glossing over: mem0 reports LLM-as-judge answer accuracy (94.4–94.8%), an end-to-end metric that includes an LLM call, while Cortyx and MemPalace report retrieval recall, a narrower, retrieval-only metric. They're related but not the same measurement. Where the comparison is closer to apples-to-apples is latency and infrastructure: Cortyx's retrieval-only activation is roughly 22ms at p95, against mem0's full pipeline at roughly 900ms–1.1s p50 because it includes an LLM call, and Cortyx requires no runtime model at all, running fully local and offline. The project ships a cortyx proof-certificate command specifically so these claims can be independently reproduced rather than taken on faith.
Where it fits
Cortyx installs with a one-line script or cargo install cortyx, indexes a project with cortyx compile, and starts an MCP server that Claude Code, Cursor, Codex, Windsurf, VS Code, and Zed can all connect to out of the box, with cortyx install auto-configuring whichever of those it detects. Beyond retrieval, it layers on a knowledge graph for temporal facts, per-agent diaries for session handoff, and a "concepts" library for sharing proven patterns across projects, all stored as the same local, git-tracked Markdown.
The trade-off is the same one that produces its speed: Cortyx delivers context, it doesn't generate answers about your code. Teams that want a system to summarize or reason over retrieved context on their behalf still need an LLM call somewhere in the loop, Cortyx just makes sure that call isn't spent re-discovering what the agent already knew last session.
Source: sorunokoe/Cortyx on GitHub.
Sage Nord