Press enter or click to view image in full size
The reproducible evaluation method I use to rank coding models on real tickets, binary CI tests, and cost per passed task.
Public coding leaderboards tell me how models perform on a shared set of problems. They do not tell me whether a model can produce correct, usable changes in my repository.
That distinction matters because every benchmark has a task distribution: the particular mix of problems, constraints, and grading rules it contains. My codebase has its own conventions, private libraries, cross-file dependencies, and recurring failure modes. A model that performs well on a public benchmark may not perform equally well on that distribution.
To answer the question I actually care about — which model works best on my codebase? — I built a private evaluation, or private eval, from 40 closed tickets. Each task uses a prompt I would realistically send to a model and a continuous integration (CI) test that returns a binary result.
The setup is intentionally small: a JSON file, a repository-specific grading function, and one portable model client. The difficult parts are not the API calls. They are choosing representative tasks and grading them without fooling yourself.
Why Public Coding Leaderboards Are Not Enough
Public benchmarks are useful as broad indicators of model capability. They answer questions such as whether a model can solve a standardized set of programming problems under a shared evaluation protocol.
My model-selection problem is narrower:
Can this model complete the kinds of tasks that appear in my repository, using patches that apply cleanly and pass the relevant tests?
A shared benchmark is unlikely to capture my codebase’s private dependencies, engineering conventions, or typical bugs. The closer an evaluation resembles my actual work, the more useful its result becomes for routing production traffic, choosing models, and comparing prompts.
That is why I treat public benchmark rankings as a starting point, not a deployment decision.
Press enter or click to view image in full size
The Five Properties of a Useful Private Eval
A private coding-model evaluation becomes useful when it has five properties.
- Real tasks: Pull tasks from closed tickets rather than synthetic puzzles. Synthetic tasks can measure general programming ability. Closed tickets test whether a model can perform work that actually appears in your repository.
- Binary grading: A task passes or it fails. For my benchmark, a successful result means the model’s patch applies and the relevant CI test passes without manual editing. Binary grading removes the ambiguity that allows wishful thinking to enter the evaluation.
- Frozen and versioned: Do not expose the expected solution to the model. Keep each version of the evaluation set unchanged while comparing models or prompts. When the codebase evolves and tasks need to be added or retired, create a new version instead of silently modifying the existing one.
- Cheap to rerun: The evaluation should be practical to run whenever a new model becomes available. If a complete sweep takes a day of manual work, it will quickly become a historical artifact rather than an active benchmark.
- Portable across models: Use one client interface for many models. Requiring a separate integration for every provider creates enough friction that you will probably test only the providers you have already integrated.
The fifth property is why I run my harness through a single OpenAI-compatible gateway. In my current setup, the same client can address more than 500 models by changing the model identifier.
The gateway does not make the benchmark valid. The task selection and grading method do that. The gateway makes the benchmark easier to rerun across providers.
The Minimal Evaluation Harness
The orchestration layer is deliberately simple:
import jsonfrom openai import OpenAI
client = OpenAI(
api_key="KEY",
base_url="https://api.cometapi.com/v1",
)
def grade(model, tasks):
passed = 0
for task in tasks:
response = client.chat.completions.create(
model=model,
messages=[
{
"role": "user",
"content": task["prompt"],
}
],
temperature=0,
)
model_output = response.choices[0].message.content
if run_test(task["test_id"], model_output):
passed += 1
return passed / len(tasks)
with open("eval_set.json", "r", encoding="utf-8") as file:
tasks = json.load(file)
models = [
"claude-sonnet-4-6",
"gpt-5",
"gemini-2.5-pro",
"qwen-3-coder",
"deepseek-v3",
]
for model in models:
score = grade(model, tasks)
print(model, round(score, 3))
That is the entire model-execution skeleton.
The difficult parts are run_test and the integrity of eval_set.json. A clean loop cannot compensate for unrepresentative tasks, leaky prompts, or subjective grading.
How I Build the Evaluation Set
I built my evaluation set from 40 tickets in my closed backlog. For each ticket, I recorded three things:
- The prompt I would realistically give to a coding model
- The identifier of the CI test that verifies the outcome
- A short note describing what makes the task difficult
Task selection is where most of the judgment lives.
I include a mix of easy one-shot functions, medium-sized bug fixes, and difficult cross-file refactors. If every task is easy, model scores cluster near the top and the benchmark cannot distinguish between them. If every task is extremely difficult, scores cluster near the bottom and the result is equally uninformative.
As a practical target, I try to build a set where my incumbent model passes approximately 60% to 80% of the tasks. That range is not a universal statistical threshold. It is simply a useful way to leave room for both improvement and regression in my evaluation.
Once the task set is selected, I freeze and version it. Comparisons are meaningful only when every candidate receives the same tasks under the same grading conditions.
Why Grading Must Be Executable
The most important grading rule is also the simplest:
Do not judge a coding-model response by whether it looks correct.
Readable explanations, plausible diffs, and confident reasoning can all accompany incorrect code. Visual inspection also tends to reward models whose answers read well, even when another model produces less polished output that actually works.
For my benchmark, grading normally follows four steps:
- Create an isolated scratch branch.
- Apply the model-generated diff.
- Run the relevant test.
- Return
TrueorFalse, then delete the branch.
The repository-specific implementation looks roughly like this:
import shutil
import subprocessdef run_test(test_id, model_diff):
branch = make_scratch_branch()
try:
apply_diff(branch, model_diff)
result = subprocess.run(
["pytest", "-q", TESTS[test_id]],
cwd=branch,
capture_output=True,
text=True,
)
return result.returncode == 0
except DiffApplyError:
return False
finally:
shutil.rmtree(branch, ignore_errors=True)
Functions such as make_scratch_branch, apply_diff, and the TESTS mapping depend on the repository. They are also where most of the initial implementation work occurs.
A patch that fails to apply receives a clean False. That is the correct result. A model has failed the task if its output cannot be applied, regardless of how convincing the surrounding explanation sounds.
Press enter or click to view image in full size
Text equivalent: For each frozen task, send the prompt to a candidate model through the same portable client. Apply the returned patch in a scratch branch. Return False if the patch cannot be applied or if the relevant CI test fails. Return True only when both checks succeed. Aggregate the results as pass rate and cost per passed task while keeping the task-set version, prompts, grading functions, and generation settings constant.
This executable grading rule is the difference between a benchmark and a horoscope.
Pass Rate Is Necessary, but It Is Not Enough
The primary quality metric is pass rate:
pass rate = passed tasks / total tasksPass rate answers a clear question: what proportion of the evaluation set did the model complete successfully?
However, pass rate alone does not capture the tradeoff between quality and price. The additional metric I find most useful is cost per passed task:
cost per passed task = total evaluation spend / passed tasksCost per passed task penalizes two kinds of weak candidates:
- Cheap models that fail so often that repeated attempts erase their price advantage
- Expensive models whose higher price is not justified by a meaningfully higher pass rate
The metric compresses the quality-versus-price tradeoff into one comparable value. A shared billing surface across models makes the calculation easier to maintain consistently.
I do not choose a model simply because it has the lowest request price. I choose it based on the cost of obtaining a result that actually passes.
How I Keep the Benchmark Honest Over Time
Model behavior can change even when the public model identifier remains the same. A model tested in June may behave differently in August because the provider has changed the backend version or serving configuration.
For that reason, I rerun the full evaluation set on a schedule and compare the results with the previous run.
When a cheaper model closes the gap with a frontier model on my tasks, I can demote the more expensive model in my router. When a model regresses, the evaluation reveals the change before it reaches users.
Adding a new candidate requires only another model identifier in the list. The minimal example above runs models sequentially, while my full runner evaluates candidates concurrently through the same gateway. A complete rerun remains a single command.
The important comparison conditions stay fixed:
- The same version of the task set
- The same prompts
- The same grading functions
- The same generation settings
Without those controls, a score change may reflect an evaluation change rather than a model change.
Common Private-Eval Mistakes
Using too few tasks
In my setup, fewer than roughly 30 tasks made the rankings too sensitive to one or two lucky passes. I now aim for 40 to 100 tasks, depending on the cost of running them.
This is an empirical rule from my own benchmark, not a universal sample-size guarantee.
Writing leaky prompts
A prompt such as “the bug is probably in the retry logic” reveals part of the solution. At that point, the evaluation is measuring the hint as well as the model.
Prompts should remain as close as possible to what I would send during normal work.
Trusting an LLM grader without verification
An LLM can be useful as a first-pass filter, but I do not treat its judgment as ground truth. LLM graders have their own preferences and can approve outputs that sound convincing but are technically wrong.
For coding tasks, deterministic tests remain the final authority. When I use model-based grading elsewhere, I manually verify a sample.
Letting the evaluation set go stale
A private eval should evolve with the codebase. Tasks that no longer represent current work should be retired, and new task types should be introduced in a new version.
An evaluation built once and never rerun is a historical snapshot, not an operational benchmark.
What the Private Eval Changed About My Workflow
The benchmark changed more than model selection.
First, I stopped arguing about models in the abstract. Questions such as “Is model X better than model Y?” now have an empirical answer based on my own tasks. That does not make the answer universal, but it makes it actionable for my repository.
Second, I started treating prompts as testable artifacts. Because the grader is binary, I can run two system prompts against the same 40 tickets and keep the one with the higher pass rate. I no longer have to choose between prompts based only on which wording feels better.
Third, I became more willing to adopt cheaper models. A lower-priced model no longer has to be accepted on faith. I can verify that it clears the required quality bar before routing real work to it.
The benchmark did not merely tell me which model ranked first. It converted a group of subjective decisions into measurements that I can defend to a skeptical teammate — or to my future self.
The Real Maintenance Cost
My private eval is lightweight because it reuses infrastructure I already have:
- A JSON file containing 40 tasks
- A
run_testfunction that wraps existing CI - A runner that sends the same task set to multiple models
In my setup, the full sweep runs with one command and completes in a few minutes of wall-clock time because the model calls run concurrently.
The recurring maintenance consists mainly of retiring tasks that are no longer representative and adding new ones as the codebase evolves. For me, that takes roughly an hour per quarter.
The main limitation is that the benchmark measures tasks with deterministic outcomes. If I cannot attach a test that returns a boolean result, I do not include the task. The evaluation therefore does not claim to measure every form of software engineering work.
The initial setup fit into a weekend because my repository already had useful CI coverage. A codebase without reliable tests would require more work at the grading layer.
The Takeaway
A public coding leaderboard tells me how a model performs on someone else’s problem distribution.
A private evaluation built from real tickets, binary CI grading, a portable model client, and cost per passed task tells me how the model performs on mine.
That difference turns “Which coding model should we use?” from an argument into a query I can rerun whenever the model landscape changes.