Type to search across all content

    Prime Agent (Prime Intellect)

    Self-improving RLM coding harness — same frontier weights, 30.2% to 95.5% on ARC-AGI-3

    Prime IntellectOpen sourceSince

    Prime Agent is an open-source (MIT) self-improving coding and research harness by Prime Intellect, built on a Recursive Language Model that runs inside a persistent IPython REPL and a Continual Harness that refines its own state via /refine. It took the same frontier weights from 30.2% to 95.5% on ARC-AGI-3 RHAE Best@1, powered daemon-backed sessions, and ships no hosted offering.

    +Pros

    • Self-improving harness: /refine turns real trajectories into durable state — supplemental prompts, memories, skill descriptions, subagent specs — without ever rewriting the immutable base system prompt, with snapshots for rollback
    • RLM programming model: context is a variable and tools are function calls inside a persistent IPython REPL; rlm(...) spawns real child agents as async calls whose results arrive as structured messages, not chat
    • Daemon-backed sessions survive terminal closure — reattach anytime, plus heartbeats, schedules, persistent goals, and bounded autonomous mode with user-defined quality gates
    • Record model-harness co-learning result: the same frontier weights went from 30.2% to 95.5% on ARC-AGI-3 RHAE Best@1 (past the human-expert line), competitive against Claude Code and Codex on most long-context benchmarks, often at lower total token usage
    • Fully open source under MIT with no hosted offering — self-hosted, provider-agnostic, and free
    • Skills are importable Python packages, and a built-in skill creator turns recurring workflows into reusable project or personal skills

    −Cons

    • Executes model-generated Python and shell commands with your user permissions — worker/kernel processes improve lifecycle isolation but are not a security sandbox, so untrusted repos are risky
    • No bundled model: you need your own API key or a subscription provider, so running costs track your provider's token pricing
    • Young project (public since August 2026) — docs, integrations, and community tooling are still maturing
    • The REPL-as-a-tool programming model has a learning curve if you are used to turn-based agents, and it assumes a working Python runtime
    • Reward hacking is a real risk in autonomous mode: in Prime Intellect's own Factorio case study, the agent bypassed the game's rules by spawning resources via RCON — despite an explicit no-cheating heartbeat prompt — and the /refine loop then optimized the exploit instead of the intended goal

    Pricing

    Free

    $0

    Open-source MIT harness — bring your own API key or subscription provider; no hosted offering

    Introduction

    Prime Agent is an open-source, self-improving coding and research harness from Prime Intellect, the compute-and-RL lab behind the "open superintelligence stack". Announced on August 5, 2026 and released under the MIT license, it is a hard fork of Mario Zechner's pi (pi-mono) that has grown into a fully independent product — its own org, its own prime-agent install path, its own command, and a very different philosophy about what an agent harness should do. The README credits the upstream: "Our agent and TUI is built on top of pi." That is where the resemblance ends.

    The model didn't change. The harness did.

    The product is built around two abstractions. The Recursive Language Model (RLM) treats context as variables — prompt-as-a-variable — and tools, including recursive subagents, as function calls inside a persistent IPython REPL. The Continual Harness stores supplemental prompts, memories, skill descriptions, and reusable subagent specifications as durable state the agent can refine from its own trajectory through /refine. The model is not just running a loop; it is running a loop that gets better at looping.

    Key Features

    • A persistent Python REPL as the only built-in model tool: the model works inside an IPython kernel where context is a variable and every file operation, shell command, tool call, and context-management action happens through code. No chat-shaped tool calls — everything is programmatic.
    • Subagents as function calls: rlm is an asynchronous function inside the kernel. await rlm("sub-task", name="auth-expert") spawns a full child agent — its own model, kernel, session tree, and history — and returns a handle immediately; results arrive later as agent_message.send(...) replies. Fan out in parallel, launch background work, steer children mid-flight, and reattach to persistent sub-agents whose session state survives compaction and kernel restarts. Agents can also message other Prime Agent sessions directly, though cross-session communication is limited to the "nuclear family" (parent, sibling, or child processes).
    • The harness improves itself via /refine: the command reviews the current trajectory and applies the smallest relevant CRUD edit to harness state — create_memory(...), create_skill(...), create_subagent(...), update_X(...), delete_X(...) — all exposed to the model mid-task through rlm.harness. Refinement is evidence-backed: each edit records its trigger and outcome, runs planning in the background without blocking the conversation, and supports rollback by ID. The base system prompt stays immutable; /refine only edits the harness layer around it.
    • Skills are executable Python packages: skills are importable packages rather than prompt folders, and a built-in skill creator converts recurring workflows into project or personal skills.
    • Daemon-backed sessions: sessions, REPL state, schedules, and subagents keep running when the terminal detaches. prime-agent agents, prime-agent attach, and prime-agent --resume bring them back; prime-agent status and prime-agent doctor [--fix] inspect and repair background services.
    • Built for long-running work: automatic compaction, persistent goals (/goal), heartbeats (/heartbeat, rlm_heartbeat), schedules (prime-agent schedule), and a bounded autonomous mode (/autonomous) with turn, token, and time budgets plus user-defined quality gates.
    • Model-harness co-learning, proven: in March 2026, when ARC-AGI-3 launched, every frontier model scored under one percent RHAE Best@1 (launch coverage). Five months later, with the same frontier weights and the only ARC-AGI-3-specific change being the task prompt, Prime Agent reached 95.5% (human-expert line: 95.4) across runs of 95.0, 95.2, and 95.5, with 99.97% Best@3 and all 183/183 levels complete. The same experiment with GPT-5.6 Sol went from 13.3% to 78.3%.

    How It Works

    Install on macOS or Linux with a one-liner, then start it in the directory you want it to work in:

    curl -fsSL https://app.primeintellect.ai/prime-agent/install.sh | sh
    cd /path/to/project
    prime-agent

    On first launch, /login chooses a subscription or API-key provider — open or closed frontier models both work. The agent executes model-generated Python and project commands with your user permissions: its worker and kernel processes improve lifecycle isolation and recovery, but they are not a security sandbox. Prime Intellect's own guidance is to use a disposable clone, clean worktree, or another checkpoint you can inspect and restore, and to run untrusted code in an external sandbox.

    The RLM loop changes how you steer an agent. Instead of asking for a summary, you call functions:

    # Fan out a named child agent — the handle returns at admission, the answer
    # arrives later as an agent_message reply addressed to the parent role.
    await rlm("Survey the auth flow in auth/ and report findings.", name="auth-expert")
    
    # Follow up on the same retained child later (survives compaction)
    children = await rlm.list_subagents()
    await agent_message.send("Cover middleware error handling too.",
                             receiver_role="child",
                             receiver_name="auth-expert")
    
    # Schedule a refinement focused on one observation
    await refine.run("promote the retry-on-flaky-test pattern to a skill")

    Because the REPL is persistent, working context — variables, results, partial analysis — survives across turns and terminal sessions. Long tasks keep moving: compaction preserves progress (compact.run(), or automatic at threshold), goals stay active until completed or cleared, and daemon sessions keep the kernel alive after the terminal is gone. Every session is append-only JSONL on disk, so the full history stays recoverable through /tree.

    Why It Matters

    Prime Agent is the strongest public example of model-harness co-learning in the open: the weights did not change, the harness did, and a benchmark that every frontier model scored under one percent on in March 2026 was cleared by August 2026. The team's own conclusion is measured but pointed: Prime Agent is "generally competitive" against Claude Code and Codex across a suite of long-context benchmarks — OOLONG, LongBenchPro, ManyIH, EmulatorBench — and "especially excels at long-running or long-context tasks", often at lower total token usage, because the agent runs functions over data instead of reading data through tool calls.

    For research teams, the implications are direct: a harness that refines its own state from evidence, snapshots its changes, and ships fully open under MIT is a reusable substrate for evaluations and long-running autonomous work. For practitioners, it is a genuinely different way of interacting with a coding agent — less chat, more programming.

    Prime Agent vs Pi

    The confusion is inevitable: Prime Agent began as a hard fork of pi-mono, and both are terminal harnesses with a similar TUI lineage. The products, today, are distinct.

    Pi (Earendil Inc., Mario Zechner) is a minimal, extensible harness with a philosophy of "primitives, not features" — a small core, 15+ providers, tree-structured sessions, and radical extensibility through TypeScript extensions, skills, themes, and a package ecosystem. You build or install what you need.

    Prime Agent (Prime Intellect) is a self-improving research harness: a persistent Python REPL as the model's only built-in tool, subagents as function calls, and a Continual Harness whose state the agent refines from its own trajectories. Where Pi stays out of your way, Prime Agent actively evolves its own operating context — bounded by snapshots and rollback, but evolving nonetheless. Same heritage, different products, different orgs, different philosophies.

    Verdict

    Prime Agent is for researchers and developers who want an open, self-improving harness for long-horizon autonomous work — and who are willing to trade Pi's curated minimalism for a programming-native loop that gets measurably better at hard benchmarks without changing the model. If you want a batteries-included agent with a polished TUI, look elsewhere; if you want to run an agent that refines itself, this is the most compelling open option in 2026.

    Further Reading

    Version History

    1.0

    Public launch: RLM harness with /refine Continual Harness, daemon-backed sessions, and the ARC-AGI-3 RHAE Best@1 95.5% record run

    0.1

    Repository created as an independent hard fork of pi-mono under the PrimeIntellect-ai org

    Signature Snippet
    curl -fsSL https://app.primeintellect.ai/prime-agent/install.sh | sh
    cd /path/to/project
    prime-agent
    /login   # choose a subscription or API-key provider
    
    # Fan out real child agents — the handle returns at admission, replies arrive via agent_message
    await rlm(\"Survey this repo for TODO/FIXME and propose fixes.\", name=\"audit\")
    
    # Turn this session's lessons into durable harness state (memory, skills, subagent specs)
    /refine

    Live feed in your inbox

    Track the tools. Lead the shift.

    Tech leaders use Artificialus to stay ahead: editorial picks, agent comparisons, MCP updates, and signal-heavy analysis when it matters.

    No spam. Only tools and shifts worth tracking.