I run agents against Amazon seller data every day at Trellis — product catalog pulls, competitor price scrapes, PPC report dumps — and every tool result one of those agents gets back lands in the context window as text the model has to read before it can do anything useful next. That is the whole case for a minimal agent harness: how much junk you let an agent carry around before you ask it to think. Earendil calls Pi 1.0 “a hardened, minimal, extensible agent harness” in its own announcement, and the number everybody is repeating is Codemode: a claimed 40% cut in prompt tokens. I want to see that hold up on a task that looks like mine before I repeat it as fact, but the mechanism behind it is the one I’d bet on.
What a minimal agent harness actually ships
Pi is Earendil’s bet that the right core for an agent loop stays small. Everything else — memory, retries, tool registries — is something you bolt on yourself instead of inheriting from a framework. MCP is native, not a plugin wrapped around an older tool-calling layer. If you already run MCP servers, Pi talks to them the way the protocol intended, instead of translating through a compatibility shim. It ships a full-screen TUI by default now, and it talks to more than 15 model providers out of the box. Swapping the model under an agent is a config change, not a rewrite.
Codemode is the part worth sitting with longer. Instead of handing the model a long list of tool schemas and asking it to emit a JSON tool call every single turn, Codemode lets the model write actual code that calls your tools directly. Only the result of running that code comes back into context. The Pi 1.0 release notes put the win at “about 40% fewer prompt tokens.” That is Earendil’s own measurement of its own feature, and nobody has published an independent comparison run on a shared task set yet. Treat it the way you’d treat any vendor’s internal benchmark: directionally believable, not something you’d bill to a client without checking yourself first.
There’s also Pi Durable, the companion package shipping alongside 1.0. It’s Earendil’s attempt at the part every agent harness eventually has to answer for: what happens to an agent’s state when the process dies mid-task. I have not run it long enough to have an opinion worth printing, and almost nobody else has either — it is days old.
The token-cost math that actually changes a decision
Here’s the arithmetic I run in my head before adopting anything like this. An Amazon seller’s product catalog pull, a competitor price scrape, a PPC report dump — each one of those tool results lands in the context window as raw text, and the model has to read all of it to do anything useful next. If Codemode really does cut 40% off the prompt-token side of that loop, and I’m running agents hundreds of times a day against seller data, that is real money, not a rounding error on an API bill. But 40% of what baseline, measured how, on what kind of task? The release notes don’t say. I’m not going to repeat the number as fact until I’ve run it myself on a task that looks like mine.
What I do believe without more proof: native MCP support is worth something regardless of the token number. Every harness I’ve touched that bolted MCP on after the fact leaked abstraction somewhere — a tool result gets formatted for the old tool-calling path, then wrapped again for MCP, double-serializing JSON that didn’t need to exist. A harness built MCP-native from day one skips that layer of translation tax on every single tool call. That tax shows up as tokens whether or not anyone benchmarks it.

You own the bugs you bolt on
Mettons you adopt Pi tomorrow. “Minimal harness you extend” also means “you now own every bug in the extension you wrote.” A batteries-included framework at least has a maintainer who already hit the dumb edge case you’re about to hit, somewhere in a GitHub issue you can search. A minimal core pushes that work back onto you.
The repo sits at 112,456 GitHub stars as I write this, and that number is a click count, not a code review — it tells you a lot of people were curious, not that 112,456 people read the state-management code and would catch a regression before it reached your production pipeline. I wrote almost this exact sentence about agents running with more trust than they’ve earned: half of that post’s budget went to checking an agent’s own output, and a thin harness with no built-in guardrails just moves more of that checking burden onto you, earlier.
Durability is where I’d watch closest. Pi Durable exists because state is the hard part of any agent harness: what happens when a long-running task gets interrupted, a tool call half-completes, or the process restarts mid-loop. Every “minimal” framework I’ve used eventually grows a durability story. The ones that grow it late, after people are already depending on the harness in production, tend to grow it badly. A brand-new package is not evidence either way on that question yet. It is a thing worth checking back on in six months, not something to build a seller pipeline on this week.
The adjacent bets worth knowing about
Two other projects showed up in the same sweep, betting on a narrower version of the same problem. context-mode is an MCP server, not a harness — it sits in front of your existing coding agent and claims a 98% cut in tool-output tokens by indexing and searching instead of dumping raw output back into context. That’s the project’s own figure too, and it ships under the Elastic License v2, which is open but not permissive the way MIT is.
openrig goes the opposite direction from Pi. It’s Apache-2.0, and it assumes you already run more than one coding agent, Claude Code and Codex side by side. It gives you a YAML file and a daemon to keep their state straight, instead of a tmux window you built by hand. Neither replaces the harness decision Pi is forcing. Both are evidence that token cost and state management are the two things every serious agent tool is converging on right now, from different directions.
Where I land
I’d try Pi on something low-stakes before I’d trust it on a Trellis pipeline that touches a seller’s ad spend. Native MCP and 15+ providers are real, checkable things today — I checked pi.dev and the provider list is there. The 40% Codemode number is a claim from Pi’s own release notes, one I’d want to reproduce on my own task before repeating it to anyone else as fact. I’d also budget real time for reading the extension code I write myself, because that code has no 112,456-star safety net under it. A minimal agent harness is a bet on how much of the hard part you want to own, not a style choice between frameworks. Pi is honest about which part it’s asking you to own. That’s worth more to me than the star count.