Benzi - 78.2% on SWE-bench Verified

10 min read Original article ↗

Benzi

A compiler backed AI coding agent

Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — and answers in O(1).
Every language runs its own tree-sitter grammar into the same query map.

github.com/oooscoos/Benzi

Most AI coding agents dump a repository into a context window and hope the model finds what matters. Benzi compiles it instead: a real compiler, built on tree-sitter, parses every file, resolves every import, builds class ancestry, and traces every identifier to its definition — one precise, queryable map, built before a single question is answered. Call flow and data flow are joined at every call site, so a bad value traces to its origin in one tool call.

Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — and answers in O(1). Every language runs its own tree-sitter grammar into that same compiled map — ten so far, plus a second engine for markup (HTML, CSS, DOM-JS).

SWE-bench Verified · 500 instances · one attempt each

78.2% 391 of 500 resolved · graded by swebench.harness.run_evaluation

The full SWE-bench Verified set — 500 real GitHub issues from twelve Python repositories — run end to end through Benzi on DeepSeek v4-flash, graded by the official SWE-bench harness inside its own per-instance Docker images. Network access to GitHub and PyPI was blocked inside every container, so nothing could look up an answer.

$37.33

total, all 500 instances

$0.095

per instance resolved

379

median lines read per instance

97%

of input tokens served from cache

27

median turns per instance

231,574

source lines read, total

Full technical report: varianttech.net/report. Every instance’s cost, tokens, turns, and lines read: varianttech.net/benchmark_swebench. The cross-harness efficiency comparison below (and the full 24-bug chart): varianttech.net/benchmark.

Try Benzi

It queries code. It writes it too.

Three ways to see it work, not screenshots. A whole app Benzi built from a single chat, a real essay on what querying a codebase actually finds, and the same agent taking apart VS Code’s own source, live. One code intelligence layer built on tree-sitter, a dedicated grammar per language, ten languages deep: Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby — plus a second engine for markup: HTML · CSS · DOM-JS.

Code writing

StallionSwipe · Python, HTML, CSS, JS

A dating app for horses, greenfielded in one chat session. Procedural SVG portraits (no image is a file — every horse is drawn in code), a swipe deck, and live AI chat where every match flirts back through a real model. Frontend, backend, and the prompts — all written by Benzi. (If the app bugs out, please open it in a new tab.)

!! interactive — try clicking around !!

Code querying

Reading DOOM’s source · C

Test code comprehension live

VS Code’s own source, resolved · TypeScript

The real repo is 1.8M lines — this indexes 923k of them: the editor core (src/vs/editor + src/vs/base), the platform services layer, and workbench’s shell/API/browser plumbing (not the 747k-line grab-bag of individual built-in features in workbench/contrib) — all TypeScript, VS Code’s own language. Built once, in just over two minutes, then cached — it updates incrementally once loaded, like on this website.

Loading. Wait time: 30 seconds.

!! interactive — try asking questions !!

Or, try any repo of your choice at all — point Benzi at any public GitHub repo and it builds the index live: varianttech.net/demo.

The architecture · one compiler, one agent loop

What's inside

Everything that falls out of actually resolving the code — from the index itself to the gates on every write.

Approach

Compiled map, not a context dump

Every file parsed, imports resolved, class ancestry built, every identifier traced to its definition — before a single question is answered. Call flow and data flow are joined at every call site, so a bad value traces to its origin in one tool call. Claude Code greps; Cursor embeds; Aider maps signatures; Benzi resolves — and answers in O(1). Every language runs its own tree-sitter grammar into that same compiled map — ten so far: Python · JavaScript · TypeScript · Java · C# · C++ · C · Go · Rust · Ruby — plus a second engine for markup: HTML · CSS · DOM-JS.

Honesty

Six states, never a guess

Every call site and every file carries one: resolved (proven in-repo edge), external (into a library, with the import evidence), candidate (ambiguous — the bounded set of possible targets, kept in full), unresolved (seen but not settled, carrying why), observed (confirmed by an actual run), unindexed (never parsed, with the reason). Whatever static analysis can’t settle is flagged as unsettled rather than guessed — and running the program is what settles it.

Safety

Blast-radius-aware, gated writes

Every write passes syntax and semantic gates against the real language parser — a broken parse auto-reverts. The model checks blast radius before it changes anything, not just after: the same analysis — the changed symbol, its callers, its holders, the selectively relevant existing tests — runs both going in and once a write lands.

24-bug cross-harness benchmark · each point is one bug

What the index actually changes

source lines read · all 24 bugs · one run each
Harness · model Lines read vs Benzi
Benzi · Sonnet 9,125—
Benzi · DeepSeek 16,4071.8×
Claude Code · Sonnet 20,7042.3×
DeepSeek Harness · DeepSeek 43,5984.8×

Lines read counts only what came back from file-read calls — grep and shell output are search, not reading. It is the one figure that means the same thing in every harness, which is why it is the one compared here.

From the 24-bug cross-harness benchmark: every harness opens more source as bugs get harder — the question is the slope. Each point is one bug, laid out easiest to hardest, left to right. Hover any point for the bug and its count.

Difficulty is Claude Code's turn count on that bug — a third-party yardstick, so no harness sets its own position on the axis. Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The figure beside each line is its slope: how many extra lines that harness opens per step of difficulty.

The same 24 bugs in the same order, with wall clock in place of lines read.

Wall clock is raw — Benzi's per-repo index build is not subtracted. Each point is that harness's most recent solved run for that bug; unsolved and unfinished runs are left out rather than plotted as fast. Benzi on Sonnet never solved http-parser and the DeepSeek harness never ran nats-server, so those two points are absent and neither enters its fit.

And the same again with dollars on the vertical axis.

Priced at the published per-token rates, same run selection as the chart above. The two DeepSeek series run an order of magnitude cheaper than the two Sonnet ones, so at this scale they sit close to the baseline. What the axis does show is the slope: Claude Code's cost climbs with difficulty faster than any other series here. Full tables behind every point live on the benchmark page.

FAQ

Does my code leave my machine?

No. In VS Code, the CLI, or MCP, the compiler and index run locally — nothing uploaded, no copy kept. Only the snippets the agent actually reads go to your model provider, same as any AI assistant, and less of them: 9,125 lines read vs Claude Code’s 20,704 on the same 24 bugs. The browser demo differs — it fetches a public repo server-side, read-only, deletes it after your session.

Do I need an API key?

No, for the browser demo. Yes for VS Code, MCP, and the CLI — run benzi-login once with your own key.

Is it actually free?

Yes, Benzi doesn’t charge. The CLI and MCP are BYOK, so you pay your own model provider. Early and in development — that’s the trade, not a paywall.

Can I point it at a private repo?

Not the web demo (public GitHub API only). Everywhere else, yes — the compiler runs locally on whatever path you give it.

How large a repo can it handle?

VS Code handles real codebases — microsoft/vscode, 923k lines, indexes in ~2 minutes, then caches. The browser demo caps at 2,000 files, 2 MB each.

How is this different from Cursor, Copilot, or Claude Code?

They search — grep or embeddings. Benzi resolves first: a real index of symbols, calls, inheritance, data flow, queried instead of guessed. Same 24 bugs, 2.3× less source read than Claude Code. Details: what the index actually changes.

How is this different from an LSP-backed MCP server?

An LSP answers at a cursor, in one open file: go-to-definition or find-references, one position and one hop at a time. Benzi compiles the whole repo up front into one index, so the questions are whole-codebase ones: transitive call trees, the path between two functions, and data flow, meaning where a bad value came from or where a return value lands. Every answer also carries a confidence tier: resolved, candidate, unresolved (with the reason), or observed. An LSP gives an answer or nothing. The runtime tracer then settles what static analysis can’t by watching what actually fires. And the compiler itself is language agnostic: ten languages run through one pipeline into one index format. An LSP setup needs a separate server per language, each installed, configured and kept running.

How is this different from CodeQL or Sourcegraph’s SCIP indexers?

They need a working build and index in batch: dependencies installed, the project compiling, a CI job measured in minutes. Benzi needs no build, works on half-finished code, and re-parses only what changed on every turn. That’s what makes gated writes possible: each edit is checked against a fresh index before the next step, not after a rebuild. One pipeline covers all ten languages instead of one indexer per language. Where that costs precision, Benzi flags the call site as candidate or unresolved instead of guessing.

How is this different from CodeGraph?

Both index instead of search, but CodeGraph retrieves — ranked candidates from a queried database. Benzi resolves — settles what a name binds to before answering, and refuses rather than guesses when a call site is ambiguous. It also models code the way an engineer reads it — file → scopes → call flow → data/control flow — not a flat symbol graph.

On CodeGraph’s own benchmark (their repos, their questions, their methodology), four models blind-judged Benzi’s MCP answers first. Full results: varianttech.net/benchmark_codegraph.

My language isn’t Python — how much do I lose?

The structural index — symbols, calls, references, inheritance, data flow — is the same across all ten languages. Only the runtime tracer is Python-only, and depth varies by language.