The scaling paradox of harness engineering
Last Updated
October 6, 2026

Spandana Kodali

Cobus Greyling
Read Time
What’s on this page
Harness engineering has expanded the surface available for improving agent performance. But it has also increased the number of dependencies that have to stay aligned for that performance to hold.
As more behavior is shaped through the harness, changes to models, tools, data, workflows or policies can ripple across the system in ways that are harder to predict and isolate.
While a better harness creates a larger optimization surface, it also makes it much harder to engineer sustained agent performance.
What is harness engineering?
Harness engineering is the practice of designing the system that surrounds an AI model so an agent performs reliably in real environments. It covers the instructions, tools, context and data access, memory, workflows, guardrails and evaluations that shape agent behavior. The model provides the intelligence while the harness determines how that intelligence is applied to a specific job.
A harness has no final configuration
In conventional software, a sufficiently tested release gives teams a relatively stable artifact to operate. Dependencies move and regressions occur, but changes can usually be traced through relatively deterministic execution paths.
Agentic systems are more tightly coupled to an environment that keeps changing.
A new model version can alter how an existing harness behaves. Changes to tools, enterprise data, workflows or policies can shift performance in ways that only become visible once the whole system is exercised again. Production adds another variable because the distribution of real requests rarely remains identical to the evaluation set assembled before launch.
The practical consequence is that the best configuration today may not remain the best configuration after the next change to the system.
This is already reflected in how agent evaluation is evolving. Anthropic recommends continuously running regression evaluations as agents change. Microsoft Foundry supports turning production traces into reusable evaluation datasets, allowing behavior observed in live traffic to shape subsequent testing.
Production therefore keeps generating new information about how well the harness is working. Maintaining performance means incorporating that information back into the system rather than treating deployment as the point at which configuration is complete.
Scale turns harness optimization into the bottleneck
For a small number of agents, this cycle is manageable manually. A problematic run can be inspected, a likely cause identified, a change made and the relevant evaluations rerun.
The workload changes quickly with scale.
An enterprise operating a growing population of agents has more configurations changing at the same time and a much larger body of behavior that has to be preserved. Improvements also become harder to judge in isolation. A change may fix the failure that prompted it while affecting performance elsewhere, so validation has to establish the impact across the broader operating envelope rather than only the cases that triggered the intervention.
This creates an optimization-throughput problem.
The constraint becomes how quickly the organization can turn a production signal into a diagnosed issue, make the appropriate change and establish with sufficient confidence that the new configuration is better overall.
The search space expands as harnesses become richer, while the validation burden expands with the number of behaviors and outcomes that must remain within acceptable bounds.
Microsoft has already made the more basic point that manual agent testing does not scale to large interaction volumes. The same pressure appears further downstream in the improvement cycle. As the number of agents and possible interventions grows, repeated human diagnosis, tuning and regression analysis becomes an increasingly significant part of the operating model.
The outcome has to become the objective function
The optimization process also needs the right target.
An evaluation suite can establish whether a change improved performance against a particular set of cases. The more important question is whether the agent became better at the job it was deployed to perform.
For a service agent, that may mean improving resolution while keeping handle time, escalation and cost within acceptable bounds. For a claims agent, faster processing only has value while decision quality and policy adherence hold. These outcomes have to be evaluated together because gains in one dimension can easily create costs elsewhere.
The target therefore has to represent the operational outcome the system is expected to produce, along with the constraints that determine whether that outcome is acceptable.
McKinsey makes a related point in its work on agentic AI economics: traditional software measures such as runtime and failure rates provide only part of the picture. Production measurement increasingly has to connect system behavior to the value and cost of the outcome being generated.
This gives harness optimization a more useful objective function. Performance can be judged against the job the agent is expected to do rather than against an isolated model metric or aggregate evaluation score.
Loop engineering: The continuous feedback loop every agent harness needs
Once outcomes become the reference point, the lifecycle starts to look different.
Production performance feeds the next evaluation cycle. When an outcome moves outside its expected range, the relevant behavior has to be isolated, the likely cause identified and an appropriate change evaluated. Validation then has to establish both that the original problem improved and that the surrounding behavior remains within bounds. The resulting configuration returns to production, where its effect becomes the next source of evidence.
The lifecycle becomes continuous:
outcomes → evaluation → diagnosis → improvement → validation → deployment → measurement
Different stages of this loop call for different engineering techniques. Deterministic checks remain appropriate where behavior can be precisely specified. Other cases benefit from model-based evaluation or reasoning. Changes with greater operational or policy impact may continue to require explicit human review.
Together, these practices are increasingly being discussed under Loop Engineering: treating agent improvement as a continuous feedback process rather than a succession of isolated development and release cycles.
For enterprise agents, the important characteristic of that loop is its connection to measurable outcomes. Production tells the system how the agent is performing against its intended job. Evaluation and diagnosis turn that signal into an actionable engineering problem. Validation establishes whether the intervention improved the system enough to return it safely to production.
Harness Engineering shapes the system around the model. Loop Engineering provides the feedback process required to keep that system aligned with its operating objectives as conditions change.
The bottleneck is moving
The first wave of agent tooling reduced much of the effort involved in creating agents. Harness engineering has since expanded the ability to shape how those agents behave in real systems.
As the number of production agents grows, more of the workload moves into maintaining their performance: interpreting production evidence, diagnosing deviations, testing improvements and establishing that the resulting configuration performs better across the outcomes that matter.
The deeper shift is in what it means for an agent to be production-ready. Readiness becomes something that has to be continually re-established as the system and its operating environment evolve. Production contributes evidence back into the engineering process, and each meaningful change creates a new obligation to prove that the agent still performs within its intended operating envelope.
That changes the lifecycle itself. The release is one point in an ongoing feedback system in which real-world performance informs the next evaluation, the next intervention and the next validation cycle.
As enterprises begin operating larger populations of AI agents, the maturity of that engineering model will increasingly depend on the quality of this feedback loop: how efficiently production evidence can be translated into validated improvement while preserving the behavior that already works.
Better harnesses will continue to expand what agents can do. Keeping those harnesses aligned with the outcomes they were built to deliver will become an equally important part of engineering them.
FAQ
What is harness engineering in AI agents?
Harness engineering is the discipline of building everything around an AI model that turns it into an agent: instructions, tools, context and data access, memory, workflows, guardrails and evaluations. Two agents running the same model can perform very differently depending on their harness, which makes it one of the main levers for improving agent performance.
Why does maintaining AI agent performance get harder at scale?
Every harness depends on models, tools, data, workflows and policies that keep changing, and a change in one area can shift behavior elsewhere. With a few agents, teams can inspect failures and rerun tests manually. With dozens or hundreds, the bottleneck becomes how quickly production signals can be diagnosed, fixed and validated without breaking behavior that already works.
What is loop engineering?
Loop engineering treats agent improvement as a continuous feedback process rather than a series of one-off releases. Production outcomes feed evaluation, evaluation drives diagnosis, and each fix is validated before redeployment, where its results become the next round of evidence. The cycle runs from outcomes to evaluation, diagnosis, improvement, validation, deployment and measurement.
How is loop engineering different from harness engineering?
Harness engineering builds the foundation: the system around the model that lets an agent do its job. Loop engineering builds on that harness, using a continuous feedback loop to automatically tune it based on production outcomes. The harness makes an agent capable; the loop keeps it delivering results as models, data, policies and real-world requests change.
How should AI agents be evaluated in production?
AI agents should be evaluated against the outcomes they were built to deliver, but outcomes alone don't explain why an agent fell short. Automated evaluation should also capture everything the agent does, including its reasoning steps, tool calls, retrieved context and decisions. Connecting that behavior to outcomes shows which harness changes will actually improve results, rather than just patching individual responses.
Share