Free
Open-source MIT harness — bring your own API key or subscription provider; no hosted offering
Self-improving RLM coding harness — same frontier weights, 30.2% to 95.5% on ARC-AGI-3
Prime Agent is an open-source (MIT) self-improving coding and research harness by Prime Intellect, built on a Recursive Language Model that runs inside a persistent IPython REPL and a Continual Harness that refines its own state via /refine. It took the same frontier weights from 30.2% to 95.5% on ARC-AGI-3 RHAE Best@1, powered daemon-backed sessions, and ships no hosted offering.
Open-source MIT harness — bring your own API key or subscription provider; no hosted offering
Prime Agent is an open-source, self-improving coding and research harness from Prime Intellect, the compute-and-RL lab behind the "open superintelligence stack". Announced on August 5, 2026 and released under the MIT license, it is a hard fork of Mario Zechner's pi (pi-mono) that has grown into a fully independent product — its own org, its own prime-agent install path, its own command, and a very different philosophy about what an agent harness should do. The README credits the upstream: "Our agent and TUI is built on top of pi." That is where the resemblance ends.
The model didn't change. The harness did.
The product is built around two abstractions. The Recursive Language Model (RLM) treats context as variables — prompt-as-a-variable — and tools, including recursive subagents, as function calls inside a persistent IPython REPL. The Continual Harness stores supplemental prompts, memories, skill descriptions, and reusable subagent specifications as durable state the agent can refine from its own trajectory through /refine. The model is not just running a loop; it is running a loop that gets better at looping.
rlm is an asynchronous function inside the kernel. await rlm("sub-task", name="auth-expert") spawns a full child agent — its own model, kernel, session tree, and history — and returns a handle immediately; results arrive later as agent_message.send(...) replies. Fan out in parallel, launch background work, steer children mid-flight, and reattach to persistent sub-agents whose session state survives compaction and kernel restarts. Agents can also message other Prime Agent sessions directly, though cross-session communication is limited to the "nuclear family" (parent, sibling, or child processes)./refine: the command reviews the current trajectory and applies the smallest relevant CRUD edit to harness state — create_memory(...), create_skill(...), create_subagent(...), update_X(...), delete_X(...) — all exposed to the model mid-task through rlm.harness. Refinement is evidence-backed: each edit records its trigger and outcome, runs planning in the background without blocking the conversation, and supports rollback by ID. The base system prompt stays immutable; /refine only edits the harness layer around it.prime-agent agents, prime-agent attach, and prime-agent --resume bring them back; prime-agent status and prime-agent doctor [--fix] inspect and repair background services./goal), heartbeats (/heartbeat, rlm_heartbeat), schedules (prime-agent schedule), and a bounded autonomous mode (/autonomous) with turn, token, and time budgets plus user-defined quality gates.Install on macOS or Linux with a one-liner, then start it in the directory you want it to work in:
curl -fsSL https://app.primeintellect.ai/prime-agent/install.sh | sh
cd /path/to/project
prime-agentOn first launch, /login chooses a subscription or API-key provider — open or closed frontier models both work. The agent executes model-generated Python and project commands with your user permissions: its worker and kernel processes improve lifecycle isolation and recovery, but they are not a security sandbox. Prime Intellect's own guidance is to use a disposable clone, clean worktree, or another checkpoint you can inspect and restore, and to run untrusted code in an external sandbox.
The RLM loop changes how you steer an agent. Instead of asking for a summary, you call functions:
# Fan out a named child agent — the handle returns at admission, the answer
# arrives later as an agent_message reply addressed to the parent role.
await rlm("Survey the auth flow in auth/ and report findings.", name="auth-expert")
# Follow up on the same retained child later (survives compaction)
children = await rlm.list_subagents()
await agent_message.send("Cover middleware error handling too.",
receiver_role="child",
receiver_name="auth-expert")
# Schedule a refinement focused on one observation
await refine.run("promote the retry-on-flaky-test pattern to a skill")Because the REPL is persistent, working context — variables, results, partial analysis — survives across turns and terminal sessions. Long tasks keep moving: compaction preserves progress (compact.run(), or automatic at threshold), goals stay active until completed or cleared, and daemon sessions keep the kernel alive after the terminal is gone. Every session is append-only JSONL on disk, so the full history stays recoverable through /tree.
Prime Agent is the strongest public example of model-harness co-learning in the open: the weights did not change, the harness did, and a benchmark that every frontier model scored under one percent on in March 2026 was cleared by August 2026. The team's own conclusion is measured but pointed: Prime Agent is "generally competitive" against Claude Code and Codex across a suite of long-context benchmarks — OOLONG, LongBenchPro, ManyIH, EmulatorBench — and "especially excels at long-running or long-context tasks", often at lower total token usage, because the agent runs functions over data instead of reading data through tool calls.
For research teams, the implications are direct: a harness that refines its own state from evidence, snapshots its changes, and ships fully open under MIT is a reusable substrate for evaluations and long-running autonomous work. For practitioners, it is a genuinely different way of interacting with a coding agent — less chat, more programming.
The confusion is inevitable: Prime Agent began as a hard fork of pi-mono, and both are terminal harnesses with a similar TUI lineage. The products, today, are distinct.
Pi (Earendil Inc., Mario Zechner) is a minimal, extensible harness with a philosophy of "primitives, not features" — a small core, 15+ providers, tree-structured sessions, and radical extensibility through TypeScript extensions, skills, themes, and a package ecosystem. You build or install what you need.
Prime Agent (Prime Intellect) is a self-improving research harness: a persistent Python REPL as the model's only built-in tool, subagents as function calls, and a Continual Harness whose state the agent refines from its own trajectories. Where Pi stays out of your way, Prime Agent actively evolves its own operating context — bounded by snapshots and rollback, but evolving nonetheless. Same heritage, different products, different orgs, different philosophies.
Prime Agent is for researchers and developers who want an open, self-improving harness for long-horizon autonomous work — and who are willing to trade Pi's curated minimalism for a programming-native loop that gets measurably better at hard benchmarks without changing the model. If you want a batteries-included agent with a polished TUI, look elsewhere; if you want to run an agent that refines itself, this is the most compelling open option in 2026.
Public launch: RLM harness with /refine Continual Harness, daemon-backed sessions, and the ARC-AGI-3 RHAE Best@1 95.5% record run
Repository created as an independent hard fork of pi-mono under the PrimeIntellect-ai org
curl -fsSL https://app.primeintellect.ai/prime-agent/install.sh | sh
cd /path/to/project
prime-agent
/login # choose a subscription or API-key provider
# Fan out real child agents — the handle returns at admission, replies arrive via agent_message
await rlm(\"Survey this repo for TODO/FIXME and propose fixes.\", name=\"audit\")
# Turn this session's lessons into durable harness state (memory, skills, subagent specs)
/refineOpenCodeReview (ocr) is Alibaba's open-source AI code-review CLI. It reads git diffs and drives an LLM agent through a hybrid architecture where deterministic pipelines handle file selection, rule matching, and comment positioning while the agent does the reasoning. Built on two years of internal use at Alibaba scale, it ships a multi-language ruleset and consumes roughly 1/9 of the tokens of general-purpose agents.
Multiplayer environment from Zed for coding with agents. Threads unify the agent conversation, the code edits, and human review in one shareable artifact, replacing the branch/PR loop. Built on DeltaDB, a CRDT-based version control layer that records every edit between commits alongside the prompts and reasoning that produced them.
Vix is a Go-native, open-source (AGPL-3.0) AI coding agent that slashes token costs by 40-50% using a stem agent architecture and Tree-sitter virtual filesystem. It rethinks the plan/execute loop — keeping LLM cache warm across Explore, Plan, and Execute phases — while shipping Programmable Workflows, Whiteboard Mode with voice AI, MCP server support, and a self-evolving agent that writes its own scheduled jobs and watchers.