CTFs Are the Only Benchmark You Can't Fake

· fantaize - systems programming & reverse engineering ·

7 min read Original article ↗

How Do You Measure an Agent?

I’ve been tracking the AI field for a while now, and it seems like every week there’s a new model that tops the benchmark charts whether it’s HLE or SimpleBench, MMLU or GPQA. It seems like these models know everything, but they struggle to apply it their knowledge and fix problems correctly. An agent should be self-aware of its own limits, for example, if it’s about to cat obfuscated_file.txt and blow up it’s context window, it should try to piece it together instead, or it should be able to learn to explore and try different strategies to solve the problem.

Agents are different from models. They shouldn’t be treated with the same restrictions as a model. Think of agents like an actual humans (like the users in the wired from Serial Experiments Lain) instead of treating them like models, Agents SHOULD have tools, like how we humans used to debug our projects on StackOverflow, or ask another human for help debugging code. Assume that all agents can use web search, and design around that instead of removing that ability from them.

So why are CTFs the only benchmark you can’t fake?

CTFs are unique from other benchmarks

Most benchmarks that test knowledge, do so in a manner where it’s just a multiple choice question, or some open-ended question that defers to a dumber model to grade it, or it uses regex which is a whole ’nother can of worms. CTFs don’t have that problem, there are random objective answers. This prevents data contamination when training LLMs, and makes it harder to benchmaxx, even when models KNOW the solution and have been trained on writeups.

Even if they have a writeup, they will still have to follow the process (tool call), and debug their problems when they eventually mess up a part of it (reasoning).

Us humans, already know that a model is probably smarter than 90% of us at math, and can spend infinite time thinking through the same problem, while we do learn considerably cheaper and faster than an AI Model, we are the cheetahs, and they are the humans with infinite stamina.

CTFs are games. They might be easier for an agent to reason through, especially if you make a stupidly easy challenge like stacking multiple different encodings on the flag itself, and so they only have to run

echo "V20xNGFGb3pjekphYWxVeVRqSlJlVTV0VlhoWmJVWnFUVlJPYWxwWFZUTk9iVlY0V1ZSR09RPT0=" | base64 -d | base64 -d | base64 -d

and instantly solve the challenge, but the harder ones are really unique, and require knowledge, connections, and tools. For example take one of my favorite CTFs (Technically this is an ARG), ae27ff.com (I know just talking about this is against the rules)

My favorite level EVER, is the brain puzzle. The user is given image of half of a brain, with the letters F S O, on it, left brain This stumps a lot of (non-technical) people. So you check inspect element, and find a hint encoded

text

42 167 150 141 164 40 151 163 40 141 40 47 155 141 147 151 143 40 156 165 155 142 145 162 
47 40 141 156 171 167 141 171 77 42 12 12 42 171 157 165 40 155 141 171 40 167 141 156 164
40 141 40 150 145 170 40 145 144 151 164 157 162 40 146 157 162 40 164 150 141 164 42 12 12

I wonder what these numbers could mean? It can’t be base64 (cause it doesn’t end with an equal sign)… Hmm.. I’m stumped.. Let’s go to dcode.fr and try to identify what it could be..

dcode cipher identifier

Seems like it’s ASCII… Let me translate so it’s human readable..

dcode ascii translator

text

"what is a 'magic number' anyway?"
  
"you may want a hex editor for that"

So you open up the hex editor on the image, and scroll around for a bit and go “Huh.. Why are there 2 image headers in the same file”, split the 2 images apart, open them side by side…

then realize the answer is: FUSION, and jump up and down because the puzzle was so beautiful and elegant.

THIS is what agents lack, THEY HAVE ALL THE KNOWLEDGE in the world, but have a hard time connecting and piecing things together! THIS is what we should be measuring. And if you think about it, we just ran through Reconnaisance, Reasoning, Tool Usage, and Vision, just for 1 challenge, this challenge didn’t require that large of a context window, nor was it particularly compute intensive (maybe for an AI model it might be), but it required application of our knowledge, searching for tools, using them properly,

What are the current problems with approach of agentic benchmarks?

A lot of them are built like one and done environments, for example that brain puzzle, will still be the same file, then and forever, but today we have the power of python, docker containers, and pseudorandom number generators. We can simulate real websites, real applications, and real games, that an agent would have to reason through, but can’t memorize; benchmaxxing would be very hard!

There’s also stuff like getting an agent to backtrack, or getting an agent to reason efficiently (I love you DeepSeek, but you reason too much), listening to audio, or forcing the agent to wait until a specific time has passed before they can continue, like we can do CRAZY stuff like time loops

Are all of these pipe dreams? Probably… Might even be a waste of tokens… But if we can train an LLM to recognize Text, Image, Video, and Audio, and serve this at Cerebras Speed? It might be possible. After that, we could begin benchmarking them by time and cost to compute, and maybe how much reasoning they did.

What if they just followed a writeup

Writeups are integral part of any CTF, which brings data contamination through memorizing the steps to achieve a task, but if you reframe the benchmark like this: “Are we measuring a model, or are we measuring an agent”. If the models themselves trained on the solutions to these challenges, wouldn’t that make them more capable? Isn’t that what we wanted in the first place, an agent that could do more in novel environments?

Currently, models are very narrow minded, given the same input of coding a website, there’s like a 90% chance it comes up with a React website, because that’s what it was mostly trained on, but if we reframe our perspective from models to agents then we can see why they’re so narrow minded.

As my good friend Richard Sutton once said, we are in the Age of Experience (I don’t fn know him btw), so let’s benchmark them like it.

Afterwords

This whole project came about when Claude Code (Opus 4.6 on the Max Plan) couldn’t solve every picoctf problem, despite there being writeups publicly online, even for the hard ones, and for some of the hard ones, it just copy and pasted the flags from them, which honestly pissed me off, cause it’s not playing the “game” correctly, so to speak, even after running 32 agents at once for 24 hours.

results after 24 hours

The Repo for that project: https://github.com/fantaize/pctf

I also read a research paper on training LLMs with CTF-Dojo with flag randomization, but I don’t think they knew what the implications were for benchmarking:

https://arxiv.org/pdf/2508.18370 (Congrats to Amazon)

This form of benchmarking (flag/answer randomization) can be to extended scientific domains, such as mathematical domains

  • Proof: (Lean, Coq, Isabelle)
  • Computation: (Randomized Coefficients)

And partially applicable to others

  • Computation Chemistry
  • A little bit of physics
  • Predicting Protein Structures (I wouldn’t fault an LLM for this one)

If you’re interested in doing ARGs, here’s a few links to help you: