Beyond Coding: The Case for Narrow AI Harnesses

· Medium ·

4 min read Original article ↗

Sumant Sogikar

Press enter or click to view image in full size

Photo by Matt Seymour on Unsplash

Harness engineering is an interesting field that has gained popularity recently. You might already be using a harness and probably don’t even realise it. Claude Code, Codex, OpenCode and numerous other coding tools are good examples of systems built around an agent harness.

Agent = Model + Harness

This equation serves as a simple way to understand harness engineering. Though there are plenty of complicated internal components that are implemented to make it professional grade, the basic idea remains the same.

So, what exactly is a harness?

Press enter or click to view image in full size

At a very high level, a harness is everything that surrounds the model and enables it to perform a task reliably.

The model itself can reason, generate text, and decide what it wants to do. But it doesn’t inherently know which tools it can use, what context it should have, what actions it is allowed to take, or how to determine whether the task was actually completed successfully.

This is where the harness comes in and it can provide the model with:

Context: the information it needs to understand the task

Tools: the ability to interact with the environment

State and memory: what has happened so far and what needs to be remembered

Permissions and policies: what the agent is allowed to do

Feedback: information about the result of an action

Verification: a way to determine whether the intended outcome was actually achieved

This creates a continuous loop around the model.

Observe → Reason → Act → Observe → Verify → Repeat

The implementation can be much more complicated, but the underlying idea remains simple. The model observes the environment, takes an action, looks at the result, and continues the loop until the task is complete.

Take a coding agent as an example.

The model is not simply asked to generate some code and return it. The harness gives it access to a repository, terminal, tests, documentation, version control and other tools. The agent makes a change, runs a test, observes the result, identifies what went wrong, makes another change and runs the test again.

The feedback loop is what makes the system considerably more capable than simply asking an LLM to write code.

Thinking about harnesses differently

Does a harness have to be for coding? Most of the examples we see today are heavily focused on software development. Software development provides a very natural feedback loop.

The same principle can be applied to other domains. Asking the right questions leads to a plethora of new ideas that can be applied to practical everyday tasks and different domains. What if the environment given to the model was Kubernetes instead of a Git repository? What if its tools were logs, metrics, events and deployment history instead of a compiler and test suite?

This leads to an idea that I find particularly interesting.

Narrow Harness

Press enter or click to view image in full size

Instead of building an increasingly general-purpose agent that can do everything, what if we build a harness around an agent that is designed to be focused on one particular job?

Imagine a harness whose entire environment is as follows:

Kubernetes + Metrics + Logs + Events + Deployment History + Runbooks

And its job is very simply to:

Observe deployment → detect abnormality → investigate → diagnose → verify

This is what I mean by a narrow harness. All the unnecessary tools, such as browsing the web, email access, managing calendars, etc., are removed from the tool set. Only the tools required for the job are made available to the agent. The agent is not becoming less capable. The environment is becoming more focused.

And this is where I think things start getting interesting.

Untapped opportunity

We have spent a lot of time making models more capable. We are also building increasingly general-purpose agents that can interact with more and more tools. What if, instead of making the agent more general, we make the environment around it more specialised?

A security harness could be built around SIEM, endpoint data and threat intelligence.

A research harness could be built around search, papers, documents and citation verification.

A finance harness could be built around transactions, accounting systems and financial rules.

A customer-support harness could be built around CRM, orders, policies and communication tools.

This makes me wonder if there is an opportunity to build an ecosystem of narrow harnesses, each designed around a specific real-world job. Instead of one agent that tries to do everything, perhaps we will have hundreds of specialised agents, each operating inside a harness designed specifically for what it needs to accomplish.

If you found this interesting, leave a comment and let me know what you think. I would love to hear your thoughts on narrow harnesses and where you see them being useful.