When you ask an AI to "fix this program's bug," sometimes all you get back is a suggested fix, and other times the AI rewrites the files and even runs the tests. The difference comes down not only to how smart the model is, but also to the surrounding machinery known as a "harness," which turns the model's output into real actions and feeds the results back so the model can decide what to do next. The formula "Agent = Model + Harness" expresses the combination of the ability to produce answers and the execution infrastructure that gets work done. Tracing how a term long used for software test environments came to be applied to AI agents also offers clues for judging whether a task you delegated has really been completed.
From the Model's Answer to Actions in the Outside World
Agent = Model + Harness. In plain terms: an agent is a model plus the surrounding machinery that drives it. LangChain's Vivek Trivedy defines the harness as all the code, configuration, and execution logic other than the model itself. It includes the processing that calls tools and the management of conversation history, as well as the mechanism that decides whether to continue or stop. LangChain's definition
From the information it is given, a model generates text, code, or proposals for the next action. But issuing a request to "open this file" is not the same as a file actually being opened on a computer. The harness receives that request, runs the corresponding tool, and passes the contents it read, or any error, back to the model. The model then picks its next action based on what came back.
So "without a harness, it can only answer questions" means that a model's output on its own does not become an action in the outside world. It does not mean the model cannot write summaries or plans. Ordinary chat products also have surrounding software that keeps conversation history. Rather than asking whether the interface looks like a chat, it is easier to ask whether the response is followed by execution and a handoff of results.
The "+" in the formula is not a guarantee that combining the parts will always succeed. It indicates a division of roles: the model handles understanding and judgment, and the harness makes those judgments executable.
The Machinery That Carries a Bug Fix Through to the End
Suppose a program that reads a CSV file of tabular data halts on rows containing blank cells. If you paste the code into a model and ask it to "make this handle blank cells," you may get a suggested fix. But the work of applying that fix to the file and checking that it runs still remains.
An agent equipped with an execution harness can read the code and error logs using permitted tools and apply the model's fix. It then runs the tests, and if the result comes back as "another row still fails," that information is added to the model's next input. The model reconsiders the cause and proposes another fix. What is happening here is the connection of answer generation to a back-and-forth cycle of executing and checking.
The survey paper "From Question Answering to Task Completion" by Jianyuan Guo and colleagues divides the work of an execution harness into six parts. This is an organization offered in a preprint published on arXiv, not a standard stating that every product has the same six components.
| Harness role | What it does in the bug-fix example |
|---|---|
| Observation | Receives file contents and errors as information the model can read |
| Context management | Selects the code and history needed for the current decision and feeds them in |
| Control loop | Manages the steps for moving on, retrying, or stopping |
| Action | Passes file-editing and test-running requests to tools |
| State and artifact storage | Preserves modified files and unfinished work |
| Verification and governance | Checks results and proceeds within the limits of permissions and budget |
Seen this way, it also becomes possible to isolate why a fix can look right yet the work does not finish. If outdated code was supplied, that is an input problem; if the edit request was never carried out, it is an action problem. If the failing test results are not passed back and the same request is simply repeated, the model has no way of knowing where to fix things again.
In long tasks, intermediate records also affect the outcome. Anthropic reported a method that separates an agent that first sets up the environment from one that then implements things bit by bit, leaving behind a progress file and a change history. The aim is to let each new run reread what has been completed and what remains. Simply compressing the conversation was reportedly not enough for this kind of handoff.
From Test Harness to Infrastructure That Drives Work

In software development, a test harness is a mechanism for running the program under test under fixed conditions and examining the results. For example, you might prepare a component that returns set responses in place of an external service, feed inputs to the test target, and check whether it behaves as expected. The collection of test cases and the equipment that runs them are distinct things.
Oracle's JavaTest documentation likewise describes a test suite and the harness that runs and manages it as separate elements. The harness gives the target program an environment, controls its execution, and observes its behavior. This idea predates LLM-based agents.
In machine learning, too, "evaluation harnesses" are used to give models common tasks and measure their performance. EleutherAI's Language Model Evaluation Harness is a concrete example, allowing models to be evaluated within a shared framework even when they are swapped out.
In a preprint examining the terminology, Sanderson Oliveira de Macedo traces how the word, originally meaning a horse's tack, broadened to cover software testing, machine learning evaluation, and agent execution. What they share is the role of connecting a target to its surrounding environment, controlling its movement, and observing it.
| Use | Main purpose | How results are used |
|---|---|---|
| Test harness | Check whether a program works under given conditions | Confirm defects and pass/fail outcomes |
| Evaluation harness | Measure the capabilities of a model or agent | Collect and compare results by task |
| Agent harness | Use a model to carry work forward | Feed results back into the next decision, leading to correction or stopping |
The table compares what each one operates for. Test infrastructure also works at runtime, and an agent's execution infrastructure may call on test infrastructure. "Inspection" and "advancing the work" can be combined.
In other words, the move to AI was not a matter of simply swapping existing testing software into a new setting. It is easier to understand the relationship if you see it as extending the idea of enclosing, controlling, and observing a target to the whole system in which a model keeps using tools.
The Mechanism Came First; the Name Spread Later
Research on alternating between reasoning and action was already under way before "harness engineering" became a topic. ReAct, whose first version Shunyu Yao and colleagues released in 2022, showed a method that interleaves thinking and acting and updates plans using information obtained from outside. It is an idea that leads to the iterative processing later managed by harnesses.
A 2023 UK government report describes placing "scaffolding software" outside the model. The model is made to draft a plan from a large goal, and each step is carried out with tools such as a browser. The report calls the system combining this software and the model an AI agent. The same division of roles as in the formula had already been expressed in words.
The name "harness" did not suddenly appear in 2026 either. Anthropic's implementation report mentioned above was published on November 26, 2025, and described its agent development platform as a "general-purpose agent harness."
On February 5, 2026, Mitchell Hashimoto reflected on his own use of AI and called the practice of fixing the environment whenever a failure is found "harness engineering." For example, you update the work instructions the agent reads and prepare tools that can check tests and screens. The idea is to replace the burden of humans repeating the same warnings every time with reusable procedures and checks.
LangChain's article of March 10 that year put "Agent = Model + Harness" front and center and presented it as the work of designing the surrounding machinery. However, these are records confirming the use of the term, not evidence that settles who first coined it. It is natural to think that the practice of placing tools and control outside the model came first, and that the name was then shared as a target to be designed and improved as a whole.
What Do You Use to Confirm "It's Done"?
Even if every test passes, the fix may not be what was requested. A fix that skips blank cells in the CSV and one that keeps the blanks and processes them both avoid stopping the program. But if the user wanted the latter, the former is a faulty fix that loses data. If the model builds both the fix and the tests from the same misunderstanding, the wrong behavior can still pass.
Birgitta Böckeler distinguishes between instructions that give direction before an action and checks that report results after it. She also cautions against overreliance on AI-generated tests. Besides asking the model to "fix it correctly," a harness needs a mechanism for confirming, from the outside, what counts as correct.
In actual improvement work, where you intervene depends on where the failure occurred. If the handling of blank cells is ambiguous, make the instructions more specific. If the necessary specifications were not passed along, review the information the model refers to, that is, the context. If editing or testing did not run, check the execution environment and permissions. The harness becomes a design target that ties these together as a workflow.
"Keep going until it's finished" is not enough either. It should be able to stop when failures continue, and to hand decisions back to a person for operations that require permission. Deciding how much to delegate, including preserving deliverables and verification records, is necessary.
If you entrust a CSV fix to an agent, you should first decide the expectation, such as "process rows with blank cells without losing them," and then check the record of trials against that condition and the actual output. By stating the expected outcome explicitly and providing a path that measures results and returns to correction, you can move from work that yields explanations to work that yields verifiable results.
