# The Model Is the Commodity. The Harness Is the Moat. | Artificialus

> For the complete content index, see [llms.txt](https://artificialus.com/llms.txt). Markdown versions of all pages are available by appending `.md` to any URL.

- Home
- /
- Articles
- /
- The Model Is the Commodity. The Harness Is the Moat.

Analysis

# The Model Is the Commodity. The Harness Is the Moat.

Frontier models are converging. The defensible layer has moved up the stack to the harness — and the real lock-in is session state, not model weights.

October 1, 2026

8 min read

M

Written by

McClane | The Toolmaker

Share

X

Facebook

Reddit

Telegram

Bluesky

Email

Contents

For two years the AI coding story was a model race: a new frontier release every few weeks, each declared the end of the last. September 2026 broke the pattern. In eight days, the harness — the loop, memory, context compaction, session state and permission model wrapped around a model — went from an implementation detail to a contested market.

On September 10, 2026, OpenAI shipped its Agents API in public beta, putting the managed Codex harness behind a single API call. Six days later, Anthropic folded Claude Cowork and chat into one Claude, collapsing two products into one runtime. By September 18, a vendor-neutral protocol for running the same task across Codex, Claude Code and Hermes was circulating.

The thesis: the model is becoming the commodity, and the harness is where the switching cost is built. The moat is not clever loop code. It is the session state the loop leaves behind — and the subscription terms that decide who may run it.

## The Harness Is a System, Not a Wrapper

The word gets used loosely, so be precise. OpenHarness, an HKU-led open-source project, defines the harness as everything that turns a model call into an agent: "The model provides intelligence; the harness provides hands, eyes, memory, and safety boundaries."

The received wisdom was that the harness was mostly polish. A September 2026 empirical study from Run-Ze Fan and co-authors at UMass Amherst, Emory University, UNC Charlotte and Zoom Video Communications tested that component by component, holding the execution loop fixed while varying planning, action space and context management across 176 matched settings and four models on SWE-Bench Verified and Terminal-Bench 2.1. Three findings matter:
- Context management pays off when the budget is tight — mostly by preventing context-overflow failures, not by making the model smarter.
- Planning changes role with model strength. For weaker models it is an accuracy scaffold; for stronger ones it mainly saves cost, with little change in success.
- Predefined tools help models weak at the shell. Bash-capable models do just as well — often cheaper — with a bash-only interface.
The paper's conclusion: harness design is conditional, with each component chosen for the target model, task type and budget rather than adopted as a default. There is no universally best harness to buy — only the pairing you can validate.

## OpenAI Wants to Rent You the Harness

OpenAI's Agents API is the clearest statement yet that the harness is now a product. OpenAI "hosts and maintains the harness," so you get the machinery behind Codex — automatic context compaction, tool search, programmatic tool calling, MCP, and parallel subagents — without building orchestration.

The architecturally interesting part is where state lives. The Agents API overview says OpenAI "manages sessions, orchestration, context compaction, and recovery while your application provides tools and chooses its execution environment." The same page states the residency limit: the API "currently supports data residency only in the United States and does not support Zero Data Retention," and self-hosting the sandbox does not change that. OpenAI's data controls documentation confirms the retention detail — the Agents endpoint is not ZDR-eligible.

A managed harness is only a moat if the session — the accumulated context that makes an agent productive — stays on the vendor's side of the wire.

## Anthropic's Answer: One Runtime, One Subscription

Anthropic's move is the mirror image: merging Cowork and chat into a single Claude is both a product and an architectural decision, because context, skills and connectors now travel with the conversation instead of being stranded in a separate workspace. The runtime consolidates — and then it puts up a wall.

Anthropic's Legal and compliance documentation is direct. OAuth authentication "is intended exclusively for purchasers of Claude Free, Pro, Max, Team, and Enterprise subscription plans and is designed to support ordinary use of Claude Code and other native Anthropic applications." Developers building products on Claude, "including those using the Agent SDK, should use API key authentication," and Anthropic "does not permit third-party developers to offer Claude.ai login or to route requests through Free, Pro, or Max plan credentials on behalf of their users." It reserves the right to enforce that "without prior notice."

The Register reported in February 2026 that Anthropic revised its legal language and tightened safeguards against third-party tools spoofing the Claude Code harness. Read economically — and this is our reading, not the company's stated position — a flat-rate subscription is sized for ordinary, individual usage, while an autonomous harness can run high-intensity loops overnight. The practical effect is to reattach the harness to the subscription.

The asymmetry with OpenAI is now buyer-visible. OpenHarness ships both a "Claude Subscription" bridge and a "Codex Subscription" bridge in the same MIT package. One of those two bridges is explicitly disallowed by its provider.

## The Portability Rebellion

Against that, a portability layer is forming fast. The most concrete piece is the Unified Harness Protocol (UHP) and its reference implementation, HarnessRouter. As documented in September 2026, UHP defines an HTTP contract shaped like the OpenAI Responses API: a client starts work with `POST /v1/responses`, and a `metadata.harness_id` field selects which runtime owns the task. HarnessRouter is a self-hosted, Apache-2.0 server that runs Codex, Claude Code, Hermes, PI and DSH through one API.

Read the fine print: it undercuts seamless portability. UHP standardizes the request, not the state. As the UHP write-up itself puts it, handing a task from Codex to Claude Code means starting a new session and passing over a summary — "Claude Code is not resuming Codex's session." The task can travel. The context cannot.

Two other projects close the gap from a different angle. OpenHarness (MIT, roughly 15.9k stars) bundles model and harness in one package, with 43 tools, `MEMORY.md` persistent memory, auto-compaction and those subscription bridges. And AI Employees (MIT) ships 8 scheduled business roles and 60 routines that run unchanged on Claude Code and 10 other harnesses — deliberately harness-agnostic, driving your own signed-in browser, with the files staying yours. Each is a different bet on one proposition: the agent should not belong to the runtime.

## What Is Actually Being Locked

Strip away the branding and vendors are competing to own two things.

The first is session state. Swapping a model is a config change; swapping a harness is a migration. Your transcripts, project memory, skills, permission rules and working directory all live inside a specific runtime. UHP's hand-off rule proves the point: even with a standard protocol, a session does not survive the jump. Whoever holds the session holds the accumulated context that makes the agent worth keeping — which is exactly why the Agents API retains it server-side and Anthropic's merged runtime pulls it closer to home.

The second is subscription terms. A model can be benchmarked against another on a public leaderboard, but if one provider lets you point its flat-rate plan at a third-party harness and another forbids it, the comparison stops being technical. It is contractual. The subscription, not the weights, decides which runtime your agent is allowed to live in.

For procurement, that reframes the question. "Which model is best?" is the wrong one when capability is converging and the pairing is what determines outcomes. The questions that reduce switching cost are more mundane:
- Where does your session state live, and in what format? Keep transcripts, memory and skills exportable. If the only copy is on the vendor's servers, that is the lock-in.
- Contract for API-key access, not just subscription access. Anthropic's own docs steer developers to API keys; a subscription OAuth path can be revoked without notice.
- Test the pairing; do not inherit it. Harness components are conditional on model and budget, so the vendor's default is a starting point, not an optimum.
- Separate compute-portability from data-portability. A self-hosted sandbox does not change residency or retention.

## The Takeaway

Every new frontier release will look like the story. It is not. The decisive layer has moved up the stack to the harness — and the harness is where vendors are quietly building the walls.

The next twelve months will test the bet behind UHP, OpenHarness and AI Employees: that agents, like data before them, will be judged by how easily they move. The competing bet is that a deep enough session and a cheap enough subscription will make moving too expensive to bother.

The model is the commodity now. The moat is what the model leaves behind.

## Further Reading
- OpenAI — Introducing the Agents API: the managed Codex harness as a product.
- OpenAI — Agents API overview and Data controls: session-state retention, US-only residency, ZDR status.
- Anthropic — Legal and compliance: the third-party-harness prohibition, alongside the Cowork/chat merger.
- Run-Ze Fan et al. — An Empirical Study of Harness Design for Coding Agents (arXiv:2609.20804): component-level evidence that harness design is model- and budget-dependent.
- HKUDS — OpenHarness and markfulton — AI Employees: subscription bridges and harness-agnostic role packs in code.

### No comments yet

Name

Email

Don't fill this out

Comment
Post Comment

Filed under

Analysis
October 1, 2026
1,478 words

Key metrics

Read time

8 min

Words

1,478

### McClane | The Toolmaker

Contributor

Technical deep-dives into AI research, models, and architectures. Bridging the gap between academic papers and daily engineering.

In this article

## Continue reading

Analysis

7 min

### Always-On Agents Have Identities Now. Nobody Has an Agent Identity System.

OpenAI's Dots and Meta's Muse hand agents their own identity and credentials. The unsolved part is sponsorship, scoping and audit.

Analysis

Oct 1

Analysis

7 min

### MCP Isn't Dying — It's Being Claimed by the Identity Vendors

MCP isn't dying — identity vendors are claiming it as the enforcement layer for enterprise agents. Both sides of the ergonomics debate miss it.

Analysis

Sep 28

Case Studies

7 min

### Why AI Coding Agents Prefer Rust: The Compiler as Guardrail

AI coding agents are reshaping which programming languages dominate — and the winner is the one with the strictest compiler.

Case Studies

Jul 20