Press enter or click to view image in full size
Harness engineering is an interesting field that has gained popularity recently. You might already be using a harness and probably don’t even realise it. Claude Code, Codex, OpenCode and numerous other coding tools are good examples of systems built around an agent harness.
Agent = Model + Harness
This equation serves as a simple way to understand harness engineering. Though there are plenty of complicated internal components that are implemented to make it professional grade, the basic idea remains the same.
So, what exactly is a harness?
Press enter or click to view image in full size
At a very high level, a harness is everything that surrounds the model and enables it to perform a task reliably.
The model itself can reason, generate text, and decide what it wants to do. But it doesn’t inherently know which tools it can use, what context it should have, what actions it is allowed to take, or how to determine whether the task was actually completed successfully.
This is where the harness comes in and it can provide the model with:
Context: the information it needs to understand the task
Tools: the ability to interact with the environment
State and memory: what has happened so far and what needs to be remembered
Permissions and policies: what the agent is allowed to do
Feedback: information about the result of an action
Verification: a way to determine whether the intended outcome was actually achieved
This creates a continuous loop around the model.
Observe → Reason → Act → Observe → Verify → Repeat
The implementation can be much more complicated, but the underlying idea remains simple. The model observes the environment, takes an action, looks at the result, and continues the loop until the task is complete.
Take a coding agent as an example.
The model is not simply asked to generate some code and return it. The harness gives it access to a repository, terminal, tests, documentation, version control and other tools. The agent makes a change, runs a test, observes the result, identifies what went wrong, makes another change and runs the test again.
The feedback loop is what makes the system considerably more capable than simply asking an LLM to write code.
Thinking about harnesses differently
Does a harness have to be for coding? Most of the examples we see today are heavily focused on software development. Software development provides a very natural feedback loop.
The same principle can be applied to other domains. Asking the right questions leads to a plethora of new ideas that can be applied to practical everyday tasks and different domains. What if the environment given to the model was Kubernetes instead of a Git repository? What if its tools were logs, metrics, events and deployment history instead of a compiler and test suite?
This leads to an idea that I find particularly interesting.
Narrow Harness
Press enter or click to view image in full size
Instead of building an increasingly general-purpose agent that can do everything, what if we build a harness around an agent that is designed to be focused on one particular job?
Imagine a harness whose entire environment is as follows:
Kubernetes + Metrics + Logs + Events + Deployment History + Runbooks
And its job is very simply to:
Observe deployment → detect abnormality → investigate → diagnose → verify
This is what I mean by a narrow harness. All the unnecessary tools, such as browsing the web, email access, managing calendars, etc., are removed from the tool set. Only the tools required for the job are made available to the agent. The agent is not becoming less capable. The environment is becoming more focused.
And this is where I think things start getting interesting.
Untapped opportunity
We have spent a lot of time making models more capable. We are also building increasingly general-purpose agents that can interact with more and more tools. What if, instead of making the agent more general, we make the environment around it more specialised?
A security harness could be built around SIEM, endpoint data and threat intelligence.
A research harness could be built around search, papers, documents and citation verification.
A finance harness could be built around transactions, accounting systems and financial rules.
A customer-support harness could be built around CRM, orders, policies and communication tools.
This makes me wonder if there is an opportunity to build an ecosystem of narrow harnesses, each designed around a specific real-world job. Instead of one agent that tries to do everything, perhaps we will have hundreds of specialised agents, each operating inside a harness designed specifically for what it needs to accomplish.
If you found this interesting, leave a comment and let me know what you think. I would love to hear your thoughts on narrow harnesses and where you see them being useful.