Benzi
A compiler backed AI coding agent
Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — and answers in O(1).
Every language runs its own tree-sitter grammar into the same query map.
Most AI coding agents dump a repository into a context window and hope the model finds what matters. Benzi compiles it instead: a real compiler, built on tree-sitter, parses every file, resolves every import, builds class ancestry, and traces every identifier to its definition — one precise, queryable map, built before a single question is answered. Call flow and data flow are joined at every call site, so a bad value traces to its origin in one tool call.
Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — and answers in O(1). Every language runs its own tree-sitter grammar into that same compiled map — ten so far, plus a second engine for markup (HTML, CSS, DOM-JS).
SWE-bench Verified · 500 instances · one attempt each
78.2%
391 of 500 resolved · graded by swebench.harness.run_evaluation
The full SWE-bench Verified set — 500 real GitHub issues from twelve Python repositories — run end to end through Benzi on DeepSeek v4-flash, graded by the official SWE-bench harness inside its own per-instance Docker images. Network access to GitHub and PyPI was blocked inside every container, so nothing could look up an answer.
$37.33
total, all 500 instances
$0.095
per instance resolved
379
median lines read per instance
97%
of input tokens served from cache
27
median turns per instance
231,574
source lines read, total
Full technical report: varianttech.net/report. Every instance’s cost, tokens, turns, and lines read: varianttech.net/benchmark_swebench. The cross-harness efficiency comparison below (and the full 24-bug chart): varianttech.net/benchmark.
Try Benzi
It queries code. It writes it too.
Three ways to see it work, not screenshots. A whole app Benzi built from a single chat, a real essay on what querying a codebase actually finds, and the same agent taking apart VS Code’s own source, live. One code intelligence layer built on tree-sitter, a dedicated grammar per language, ten languages deep: Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby — plus a second engine for markup: HTML · CSS · DOM-JS.
Code writing
StallionSwipe · Python, HTML, CSS, JS
A dating app for horses, greenfielded in one chat session. Procedural SVG portraits (no image is a file — every horse is drawn in code), a swipe deck, and live AI chat where every match flirts back through a real model. Frontend, backend, and the prompts — all written by Benzi. (If the app bugs out, please open it in a new tab.)
!! interactive — try clicking around !!
Code querying
Reading DOOM’s source · C
Test code comprehension live
VS Code’s own source, resolved · TypeScript
The real repo is 1.8M lines — this indexes 923k of them: the editor core
(src/vs/editor + src/vs/base), the platform services layer, and workbench’s
shell/API/browser plumbing (not the 747k-line grab-bag of individual built-in features in
workbench/contrib) — all TypeScript, VS Code’s own language. Built once, in
just over two minutes, then cached — it updates incrementally once loaded, like on this website.
Loading. Wait time: 30 seconds.
!! interactive — try asking questions !!
Or, try any repo of your choice at all — point Benzi at any public GitHub repo and it builds the index live: varianttech.net/demo.
The architecture · one compiler, one agent loop
What's inside
Everything that falls out of actually resolving the code — from the index itself to the gates on every write.
Approach
Compiled map, not a context dump
Every file parsed, imports resolved, class ancestry built, every identifier traced to its definition — before a single question is answered. Call flow and data flow are joined at every call site, so a bad value traces to its origin in one tool call. Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — and answers in O(1). Every language runs its own tree-sitter grammar into that same compiled map — ten so far: Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby — plus a second engine for markup: HTML · CSS · DOM-JS.
Honesty
Six states, never a guess
Every call site and every file carries one: resolved (proven in-repo edge), external (into a library, with the import evidence), candidate (ambiguous — the bounded set of possible targets, kept in full), unresolved (seen but not settled, carrying why), observed (confirmed by an actual run), unindexed (never parsed, with the reason). Whatever static analysis can’t settle is flagged as unsettled rather than guessed — and running the program is what settles it.
Safety
Blast-radius-aware, gated writes
Every write passes syntax and semantic gates against the real language parser — a broken parse auto-reverts. The model checks blast radius before it changes anything, not just after: the same analysis — the changed symbol, its callers, its holders, the selectively relevant existing tests — runs both going in and once a write lands.
24-bug cross-harness benchmark · each point is one bug
What the index actually changes
| Harness · model | Lines read | vs Benzi |
|---|---|---|
| Benzi · Sonnet | 9,125 | — |
| Benzi · DeepSeek | 16,407 | 1.8× |
| Claude Code · Sonnet | 20,704 | 2.3× |
| DeepSeek Harness · DeepSeek | 43,598 | 4.8× |
Lines read counts only what came back from file-read calls — grep and shell output are search, not reading. It is the one figure that means the same thing in every harness, which is why it is the one compared here.
From the 24-bug cross-harness benchmark: every harness opens more source as bugs get harder — the question is the slope. Each point is one bug, laid out easiest to hardest, left to right. Hover any point for the bug and its count.
Difficulty is Claude Code's turn count on that bug — a third-party yardstick, so no harness sets its own position on the axis. Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The figure beside each line is its slope: how many extra lines that harness opens per step of difficulty.
The same 24 bugs in the same order, with wall clock in place of lines read.
Wall clock is raw — Benzi's per-repo index build is not subtracted. Each point is that harness's most recent solved run for that bug; unsolved and unfinished runs are left out rather than plotted as fast. Benzi on Sonnet never solved http-parser and the DeepSeek harness never ran nats-server, so those two points are absent and neither enters its fit.
And the same again with dollars on the vertical axis.
Priced at the published per-token rates, same run selection as the chart above. The two DeepSeek series run an order of magnitude cheaper than the two Sonnet ones, so at this scale they sit close to the baseline. What the axis does show is the slope: Claude Code's cost climbs with difficulty faster than any other series here. Full tables behind every point live on the benchmark page.
FAQ
Does my code leave my machine?
No. In VS Code, the CLI, or MCP, the compiler and index run locally — nothing uploaded, no copy kept. Only the snippets the agent actually reads go to your model provider, same as any AI assistant, and less of them: 9,125 lines read vs Claude Code’s 20,704 on the same 24 bugs. The browser demo differs — it fetches a public repo server-side, read-only, deletes it after your session.
Do I need an API key?
No, for the browser demo. Yes for VS Code, MCP, and the CLI — run
benzi-login once with your own key.
Is it actually free?
Yes, Benzi doesn’t charge. The CLI and MCP are BYOK, so you pay your own model provider. Early and in development — that’s the trade, not a paywall.
Can I point it at a private repo?
Not the web demo (public GitHub API only). Everywhere else, yes — the compiler runs locally on whatever path you give it.
How large a repo can it handle?
VS Code handles real codebases — microsoft/vscode, 923k lines, indexes in
~2 minutes, then caches. The browser demo caps at 2,000 files, 2 MB each.
How is this different from Cursor, Copilot, or Claude Code?
They search — grep or embeddings. Benzi resolves first: a real index of symbols, calls, inheritance, data flow, queried instead of guessed. Same 24 bugs, 2.3× less source read than Claude Code. Details: what the index actually changes.
How is this different from an LSP-backed MCP server?
An LSP answers at a cursor, in one open file: go-to-definition or find-references, one position and one hop at a time. Benzi compiles the whole repo up front into one index, so the questions are whole-codebase ones: transitive call trees, the path between two functions, and data flow, meaning where a bad value came from or where a return value lands. Every answer also carries a confidence tier: resolved, candidate, unresolved (with the reason), or observed. An LSP gives an answer or nothing. The runtime tracer then settles what static analysis can’t by watching what actually fires. And the compiler itself is language agnostic: ten languages run through one pipeline into one index format. An LSP setup needs a separate server per language, each installed, configured and kept running.
How is this different from CodeQL or Sourcegraph’s SCIP indexers?
They need a working build and index in batch: dependencies installed, the project compiling, a CI job measured in minutes. Benzi needs no build, works on half-finished code, and re-parses only what changed on every turn. That’s what makes gated writes possible: each edit is checked against a fresh index before the next step, not after a rebuild. One pipeline covers all ten languages instead of one indexer per language. Where that costs precision, Benzi flags the call site as candidate or unresolved instead of guessing.
How is this different from CodeGraph?
Both index instead of search, but CodeGraph retrieves — ranked candidates from a queried database. Benzi resolves — settles what a name binds to before answering, and refuses rather than guesses when a call site is ambiguous. It also models code the way an engineer reads it — file → scopes → call flow → data/control flow — not a flat symbol graph.
On CodeGraph’s own benchmark (their repos, their questions, their methodology), four models blind-judged Benzi’s MCP answers first. Full results: varianttech.net/benchmark_codegraph.
My language isn’t Python — how much do I lose?
The structural index — symbols, calls, references, inheritance, data flow — is the same across all ten languages. Only the runtime tracer is Python-only, and depth varies by language.
Get started
Source, README and issues on GitHub: github.com/oooscoos/Benzi
Benzi is completely free to use.
- In the browser — paste any public GitHub repo at varianttech.net/demo; no install, no signup. Read-only: ask it questions, explore the map, nothing writes to the repo. This is the demo — click here to see what it can do.
- In VS Code — the same compiler, but with edit access: chat, graph, and Benzi actually writing code in your own project. VS Code Marketplace. This is the real tool — click here to use it.
- MCP — the same compiled index, exposed as tools over MCP for whatever agent you
already run: Claude Code, Cursor, or your own harness.
pip install benzi, then point your MCP client atbenzi-mcp. This is the benzi index without the agentic loop — output quality will depend on your agent/harness. - Headless — the same agent as VS Code, from your own terminal:
pip install benzi, thenbenzi-headless <repo> "your question". This is Benzi for scripts and CI — no editor needed. (pypi.org/project/benzi)
Run benzi-login once to authenticate before using the VS Code extension, MCP, or
headless — the same command lets you update your model or key again later too.