# Idyllic Labs — complete text > Idyllic Labs is an applied R&D lab. It builds tools for human superagency: structures that turn machine intelligence into causal power for people and organizations. Index of pages: https://idylliclabs.com/llms.txt · Sitemap: https://idylliclabs.com/sitemap.xml --- # Idyllic Labs URL: https://idylliclabs.com APPLIED R&D LAB # Tools for human superagency. We believe artificial intelligence is a transformational technology because it makes intelligence a resource, something that can be acquired and directed. Agency is causal power, the ability to make things happen in the world, and people and organizations expand it by building structures that act on their behalf. Superagency is agency expanded with machine intelligence, and Idyllic Labs builds tools for it: tools that shorten the distance between wanting something to exist and it existing, and that bring out potential in people that would otherwise stay dormant. We use everything we build ourselves first, and we publish what we learn. PROJECTS - [PLATFORM — Busibody — Infrastructure for businesses operated by AI agents. It handles incorporation, banking, payments, and email, and the owner approves anything consequential. Each business is a more granular startup, cheap enough to run that a niche too small or too short-lived to justify a person’s time can still be worth operating.](https://busibody.ai) - [BOOK — Elements of Agentic System Design — A conceptual framework for the design of intelligent systems from first principles. Ten elements — from context to learning — that decompose any agentic system into the mechanisms that implement it. An open repository and a forthcoming book.](https://github.com/idyllic-labs/elements-of-agentic-system-design) - [PLANNED — Convergence — A desktop harness for directing many coding agents. Context made visible, parallel work shown as branches converging on a trunk, and the operator’s judgment applied in small comparisons that lock as they are accepted.](https://idylliclabs.com/projects/convergence) - [PLANNED — Editorial — A system for writing with AI that converges. Intent locked above prose, candidates in small judgable units, sections that freeze when accepted, and every verdict compiled into editorial law. Built on Convergence.](https://idylliclabs.com/projects/editorial) WRITING - [2026 ESSAY — Intelligence is water — Intelligence behaves like water. It presses on every surface it touches, passes through whatever opening exists, and takes the shape of its container, so the container sets the quality of what comes out.](https://idylliclabs.com/writing/intelligence-is-water) - [2026 ESSAY — The agent body — Agent products all run on the same models, tools, and prompts, so the parts cannot explain why some improve and hold users while others stay demos. What separates them is the body: what persists between calls, where the boundary sits, and why the boundary is worth having.](https://idylliclabs.com/writing/the-agent-body) - [2026 ESSAY — The granular startup — Steve Blank called a startup a temporary organization searching for a repeatable business model, and nothing in the definition says how many people the search takes. When agents supply the labor the set of markets worth searching has to be recomputed, and at the limit the exploration of the markets startups ignore can itself be an automated process.](https://idylliclabs.com/writing/the-granular-startup) [All writing →](https://idylliclabs.com/writing) --- # Superagency URL: https://idylliclabs.com/superagency # Superagency Superagency is the expansion of human causal power through intelligent structures. Idyllic Labs builds tools for it, for individuals and for organizations. Agency is causal power. A person with more agency can cause more of what they intend: they can make a product exist, reach the people it concerns, change how something is done in their trade. Humans have always expanded their causal power by building structures outside themselves that they control. Writing, money, firms, and software are all structures of this kind, and each one carries a person's intent further than their own hands and hours could carry it. Organizations sit under the same limit at a larger scale: what a company can attempt is bounded by its headcount and its coordination costs, and the bound loosens the same way, by building structures. What is new is that the structures can now think. ## The intelligence revolution Intelligence has become a resource. Judgment, planning, research, and writing used to be obtainable only from people, at the pace of human attention. A machine now supplies them, metered like electricity, and the question for a builder changes accordingly, from who to hire to what structure to pour the resource into. A large language model does two things well. It can generate many candidate plans, drafts, and approaches from one starting point, which is divergence, and it can judge those candidates against a goal and narrow them to the ones that work, which is convergence. A system that can do both can search, and search power is what greater intelligence amounts to in practice: more of the space of possible solutions examined per unit of time. **FIG. 1 · SEARCH AS DIVERGENCE AND CONVERGENCE** *Each branch is a candidate approach. Judgment narrows the search to the path that best satisfies the intent.* An AI agent runs this search in a loop. It tries an approach, looks at what happened, and tries again with what it learned. For a person, trying again is the expensive part — the fortieth attempt at a stubborn problem costs attention, morale, and hours that the first attempt did not. For a machine, the fortieth attempt costs the same as the first. Determination used to be a rare trait of people. It is now priced in compute: as much pathfinding as you are willing to pay for. An engineer recently wrote about being surprised by his own agent. Someone sent the agent a voice memo on Telegram, and within seconds it had found a transcription tool it was never told about, read the message, and replied in generated voice. Nothing in its instructions covered voice memos. It searched what was available, found an opening, and converged on an answer, and that is not a surprising outcome, because pathfinding toward whatever is possible is exactly what the loop is built to do. ## The two camps There are two readings of this technology. In the first, machine intelligence is a substitute for human effort: whatever a person can do, a machine will eventually do cheaper, so human work loses its point and human skill loses its value. We take this reading seriously. Substitution is real, and some of it will hurt. But the reading treats people as suppliers of labor and nothing else, and that is where it goes wrong. We hold the second reading. Most of what people could contribute never gets expressed, and the missing ingredient was never judgment. Plenty of people can tell a good product from a bad one, know exactly what their trade needs, or have an accurate picture of what their customers want. What they lack is labor: the founder needs engineers, the shop owner needs staff, the writer needs an editor and a publicist. When intelligent structures supply the labor, the judgment that already existed gets expressed, and things come to exist that otherwise would not have. That is the position we build from: this technology is for the liberation of human potential, and machines are its instruments, not its replacement. ## What an agent is Humans have always pushed computation out of their heads and into structures. A ledger computes a balance, a procedure computes a decision, an org chart computes who handles what. An AI agent is one more externalized computational structure: a piece of software that uses an AI model to do work on its own. It can write code, send email, search the web, and operate other software, and it keeps working while the person it works for is doing something else. We describe an agent as having a mind and a body. The mind is the model and the context it is given: it does the thinking, it holds no state of its own, and one mind can be swapped for another. The body is everything that persists between runs: the accounts, the records, the obligations, the accumulated rules, the identity the work belongs to. Most of the field's attention goes to the mind, to better models and better prompts. We put ours on the body, because the body is what turns many runs of an interchangeable mind into one enduring thing in the world, and it is the part of the structure the person actually owns. ## How we work We build tools and use them ourselves before offering them to anyone else. The lab runs real businesses on its own software, so we meet each tool's problems before a customer does. Running them for real also does something no simulation can: a search loop only converges when something pushes back on it, and real customers, real payments, and real replies are what push back. An agent business that runs without that contact produces slop. One that runs with it gets corrected every day. The agents do not act freely. Consequential actions, such as spending money, publishing in a business's name, or making a commitment to a customer, wait for a person's approval, and when experience shows that a kind of action reliably goes well, we move it under a standing policy and the person reviews results instead of approving each one. When an agent's work comes back wrong, we write down what was wrong and why, and the note becomes a rule that every future attempt starts from. A mistake corrected this way tends to stay corrected, because the correction lives in writing rather than in anyone's memory. We publish what we learn as we go. ## What we build toward We want to minimize the distance between a person wanting something to exist and that thing existing. Today the distance is measured in years of a person's life, or in the cost of hiring and managing other people, and most ideas never cross it. The ideas that go untried are usually not considered and found bad. There is simply no time to consider them. The same distance hides value in markets. Many small businesses are never started because a person's time cannot be justified: the niche is too small, the margin too thin, or the opportunity too short-lived to repay years of attention. That value seeps through the cracks. A structure that carries the effort changes the arithmetic, and a niche that will only exist for three months can still be worth serving when serving it costs a person an hour of review a week. The larger store of dormant value is in people. Most of what a person knows, wants, and can judge never becomes anything, because becoming something took labor they did not have. We build the structures that let it become something, and everything we build is a step toward that. --- # The Lab URL: https://idylliclabs.com/lab THE LAB # The lab. MODEL ## What the lab does Idyllic Labs is an applied research lab. Intelligence has become a resource: anyone can buy it by the token, the way they buy computing power or storage. On its own it does nothing, so we build the structures that turn it into causal power for people and organizations — tools to build software without a team, and to start and run things that would otherwise take a company. We use everything we build ourselves first: the lab runs small businesses operated by AI agents on its own tools, and what we learn from running them decides what we build next. The lab's public work is [Busibody](https://busibody.ai), infrastructure for businesses operated by AI agents, and the essays and lab notes at [/writing](https://idylliclabs.com/writing). | FIG. 1 · SITE VITALS, BUILD INDEX | | | --- | --- | | INDEXED FROM SOURCE | | | 18 | PAGES | | 12 | ESSAYS | | 20,861 | WORDS | *Page count expands the writing route once per published essay. Words include essay titles, summaries, and bodies in this build.* LINEAGE ## Where the lab sits The lab works in the tradition of Bell Labs, Xerox PARC, and Ink & Switch: institutions where the people who designed a system also operated it, so the research stayed in contact with practice. Idyllic Labs works the same way. We publish an essay after the practice it describes has been used in one of the businesses we run. The idea the lab is organized around, human superagency, is described in [its own essay](https://idylliclabs.com/superagency). How the lab thinks about business descends from Steve Blank, who called a startup a temporary structure that searches for a working business through experiments, and from Paul Graham, who called it a mosquito: small, lean, and built to do one thing. The businesses the lab runs carry that idea further down in scale, with agents doing most of the searching. **FIG. 2 · APPLIED R&D LINEAGE** *Each trace begins at an institution where research and operation shared a feedback loop, then converges on the lab's present practice.* CORRESPONDENCE ## How to reach the lab Mail to the lab goes to [william@idylliclabs.com](mailto:william@idylliclabs.com), and a person reads it. Corrections and disagreements are welcome, and the useful ones are cited in the writing. COLOPHON ## How this site is made | FIG. 3 · COLOPHON | | | --- | --- | | TYPE | Plus Jakarta Sans for headings, Geist at 18px for body text, and Berkeley Mono for labels, tags, and figures. | | SYSTEM | One set of design tokens: a near-black background, a single content panel, hairline rules, a light text ramp, and no shadows. | | FIGURES | Labeled boxes with a hairline border, a mono header, and a footnote, like the one around this list. | | BUILD | Next.js, deployed from commit 67954f2. | *The typefaces, design tokens, and build of idylliclabs.com.* --- # Convergence URL: https://idylliclabs.com/projects/convergence # Convergence Convergence is a desktop harness for directing many coding agents at once. It is in design. This page records the problem it addresses and the intent that governs how it will be built. ## The problem Working with many agents happens today through interfaces built for one conversation. The operator's entire view of a fleet is a chat stream. What context each agent was given, and what it currently holds, cannot be seen or steered. Parallel branches of work exist only as text scrolling past, so the state of the whole — what is exploring, what is finished, what is waiting on a decision — lives nowhere but the operator's head. The scarcest input in such a system is the operator's judgment, and it has no tooling at all. There is no way to compare two candidate results side by side, no way to accept one part of the work and freeze it, no way to flag one line and leave the rest untouched. Every judgment is typed into the same chat box as everything else, and nothing the operator decides is kept. ## The macrostate One requirement governs the design: the operator's input lands only at the points where it scales — stating what must be true before the work, and judging what comes back after it — and both are made cheap. Concretely: Context flow is visible: what each agent was given, what it holds, when it changed. Parallel work appears as branches converging on a trunk through a merge queue, because progress accrues only at the merge. Candidate results arrive in units small enough to judge in seconds, side by side; judgment cost sizes every surface. Accepted work locks, a flag produces an in-place patch, and the work only moves closer to done. And every judgment the operator makes is captured and encoded — a verdict becomes a rule, a rule becomes a check — so taste is paid for once and governs every run that follows. ## Resolution with the philosophy [Superagency](https://idylliclabs.com/superagency) is causal power through structures, and attention is the input that does not scale. [How to work with many agents](https://idylliclabs.com/writing/working-with-many-agents) develops the mental models: macrostates over microstates for the briefing side, convergence to one trunk for the judging side. [Shipping is merging](https://idylliclabs.com/writing/shipping-is-merging) locates the control point of agent work at the join, and [Encoding judgment as rules](https://idylliclabs.com/writing/encoding-judgment-as-rules) shows how a person's decisions govern runs they never see. Convergence is those claims built as one instrument. Agent engines remain the engine underneath; Convergence is the harness and the visible surface. Status: defined and queued. The first prototype runs as skills and scripts inside the agent CLI, with the interface mocked through MCP; a dedicated application exists for the one capability a hosted assistant cannot provide, which is programmatic control of third-party coding agents. --- # Editorial URL: https://idylliclabs.com/projects/editorial # Editorial Editorial is a system for writing with AI, built on the same machinery as [Convergence](https://idylliclabs.com/projects/convergence) with prose as the material. It is in design. This page records the problem and the intent. ## The problem Writing with AI does not converge. Feedback on a draft triggers a rewrite, the rewrite arrives in a new uniform register, and the next round of feedback produces a different essay rather than a closer one. The failure has mechanics. A model writing in one pass locks into a register and reinforces it, so the whole draft gets exactly one texture. Matching a corpus reproduces the average of the corpus and loses the variance where a voice actually lives. And rules compiled from an author's reactions can only forbid; a style is a positive object, and no list of bans generates one. Meanwhile the judgments the author gives — the only signal that matters — evaporate with the conversation that carried them, and the same corrections are paid for again on every draft. ## The macrostate An essay converges the way software does: layer by layer, with locks. Intent sits above prose — the audience, what they believe now, what they leave with — and is decided before sentences exist, so an edit routes to the layer it belongs to instead of being wordsmithed downstream. Candidates arrive in small units, a paragraph or an opening, because prose is expensive to compare and judgment cost sizes the units. Accepted sections freeze; a flagged sentence produces an in-place patch; the draft only moves closer. Every verdict compiles into editorial law: a style guide with worked examples, machine-checkable rules, a library of accepted work. An editor pipeline applies the law before anything reaches the author, so what the author reads is already filtered, and what the author decides is never spent twice. ## Resolution with the philosophy Editorial is [Encoding judgment as rules](https://idylliclabs.com/writing/encoding-judgment-as-rules) and [Shipping is merging](https://idylliclabs.com/writing/shipping-is-merging) applied to prose. It exists because the lab needs it: the essays on this site are produced by this process run by hand, and the tool is that process made into an instrument. Its first user is the lab's own corpus. It is built on Convergence because convergence machinery does not care what the candidates are; prose is the second material. Status: defined and queued behind Convergence, prototyped the same way — skills and scripts inside the agent CLI, the interface mocked through MCP. --- # Writing URL: https://idylliclabs.com/writing WRITING 12 PIECES · 05 SHELVES # What we learn building and operating our tools. READING ORDER 01—12 SHELF 01 · 03 PIECES ## Persistent state An agent system rents its model and owns everything else: the record of what it has done, the judgments on its past work, the rules and checks that encode its craft, and its standing with the people it serves. The owned parts are the ones that accumulate value, and the model is the part to plan on replacing. - [07 — →08 — What compounds in an agent system — In a blind comparison, a mid-tier model matched a frontier model on every dimension our tooling had encoded. The model is the part of an agent system to plan on replacing, and what compounds is the owned layer: the history, the verdicts, the encoded craft, and the standing it earns. — ESSAY — JUL 2026 — 5 MIN](https://idylliclabs.com/writing/what-compounds-in-an-agent-system) - [08 — →09 — The value of the harness — A model in a chat window gives you drafts you still have to check. The same model inside a real pipeline produces finished work. This essay is about the harness, everything built around the model, and why each new model release multiplies whatever you have built there. — ESSAY — JUL 2026 — 6 MIN](https://idylliclabs.com/writing/systems-appreciate-prompts-depreciate) - [11 — →12 — Ideas with bodies — An idea held in a head or a note cannot be found by a stranger, cannot take a payment, and cannot learn from silence. Agents have made it cheap to form the body an idea needs to meet its market, meaning an address, an inbox, an account, and records that persist. — ESSAY — JUL 2026 — 6 MIN](https://idylliclabs.com/writing/ideas-with-bodies) SHELF 02 · 01 PIECES ## Measurement Most of what agents send into the world gets no response. When each stage of the path from sending to reply carries its own measurement, a non-response identifies the stage that failed, and each round of sends becomes an experiment that improves the next one. - [09 — →10 — Silence is a signal — Most work an agent sends into the world gets no reply, and a bare no-reply teaches the system nothing. Splitting the outcome into stages, received, opened, clicked, answered, turns silence into an error signal the system can actually learn from. — ESSAY — JUL 2026 — 5 MIN](https://idylliclabs.com/writing/instrumenting-agent-work) SHELF 03 · 01 PIECES ## Governance The humans behind an agent system cannot review thousands of runs directly. Corrections recorded as written rules, and enforced by automated checks, carry their judgment into runs they never see, and the human role narrows to approval and correction. - [10 — →11 — Encoding judgment as rules — Most of a person's judgment can be written down, and the written form is what governs. Corrections recorded as rules reach thousands of runs the person never sees, and every review makes the rules more complete. — ESSAY — JUL 2026 — 6 MIN](https://idylliclabs.com/writing/encoding-judgment-as-rules) SHELF 04 · 06 PIECES ## Foundations Intelligence is now a resource anyone can buy, and a bought resource does nothing until it has a structure to run through. An agent is one such structure, an externalized piece of computation with a mind and a body, and agency is what the structures exist to expand: the power of a person or an organization to cause what they intend. - [01 — →02 — Intelligence is water — Intelligence behaves like water. It presses on every surface it touches, passes through whatever opening exists, and takes the shape of its container, so the container sets the quality of what comes out. — ESSAY — JUL 2026 — 5 MIN](https://idylliclabs.com/writing/intelligence-is-water) - [02 — →03 — Search that does not tire — An agent loop is a search algorithm: propose an action, carry it out, read the result, adjust. Retrying is the part of search that exhausts a person, and it is the part a machine gets at the price of a model call. — ESSAY — JUL 2026 — 6 MIN](https://idylliclabs.com/writing/search-that-does-not-tire) - [03 — →04 — Shipping is merging — Agents made it nearly free to open parallel branches of work. A branch becomes progress only when its result is merged back into the one thing being built, and the merge is the step that did not get cheaper. — ESSAY — JUL 2026 — 6 MIN](https://idylliclabs.com/writing/shipping-is-merging) - [04 — →05 — How to work with many agents — The mental models that scale one person from ten agents to a thousand: hold macrostates, not microstates, and converge everything into one trunk. — ESSAY — JUL 2026 — 7 MIN](https://idylliclabs.com/writing/working-with-many-agents) - [05 — →06 — The agent body — Agent products all run on the same models, tools, and prompts, so the parts cannot explain why some improve and hold users while others stay demos. What separates them is the body: what persists between calls, where the boundary sits, and why the boundary is worth having. — ESSAY — JUL 2026 — 7 MIN](https://idylliclabs.com/writing/the-agent-body) - [06 — →07 — Nothing grown in isolation survives transplant — Shipping speed and growth location look like one variable and are two: a team can launch every week while everything it builds grows against its own model of the market. Readiness is a product of contact with the destination, not a precondition for it. — ESSAY — JUL 2026 — 6 MIN](https://idylliclabs.com/writing/nothing-grown-in-isolation-survives-transplant) SHELF 05 · 01 PIECES ## Economics A startup is a temporary structure built to search for a working business, and the search has always been priced in human time. When agents supply the labor, the search is priced in compute instead, and markets too small, too thin, or too short-lived to repay a person's attention become worth searching. - [12 — →01 — The granular startup — Steve Blank called a startup a temporary organization searching for a repeatable business model, and nothing in the definition says how many people the search takes. When agents supply the labor the set of markets worth searching has to be recomputed, and at the limit the exploration of the markets startups ignore can itself be an automated process. — ESSAY — JUL 2026 — 5 MIN](https://idylliclabs.com/writing/the-granular-startup) END OF INDEX · CONTINUE AT 01 [WILLIAM@IDYLLICLABS.COM](mailto:william@idylliclabs.com) · [RSS](https://idylliclabs.com/writing/rss.xml) --- # Intelligence is water URL: https://idylliclabs.com/writing/intelligence-is-water Area: Foundations · Published: 2026-07-17 Intelligence behaves like water. It presses on every surface it touches, passes through whatever opening exists, and takes the shape of its container, so the container sets the quality of what comes out. Peter Steinberger, the developer behind the open-source agent OpenClaw, tells a story about sending his own agent a voice memo. Within seconds the agent identified the unknown audio file, found ffmpeg on the local disk, found an OpenAI key in its environment, transcribed the memo, and replied with the transcript, none of which he had programmed.[1] He kept introducing the agent as magical because of moments like this one. The usual description of intelligence is a planner. An intelligent system, in that picture, understands a goal, forms a strategy, and executes the steps, and a better system is one with a better plan. Applied to AI agents, the picture produces a specific expectation: an agent is a worker, you brief it, and it carries out the brief. When an agent then does something nobody briefed it to do, the planner picture makes the behavior look like initiative. Initiative from a machine is unsettling. We describe intelligence as water because the comparison predicts things the usual description misses. Water has no plan and no goal. It flows downhill, presses on every surface it touches, and passes through the first crack that yields, and whether it floods a basement or turns a turbine depends entirely on what contains it. We think intelligence behaves the same way, in machines and in people, and most of our practical decisions about building with AI follow from taking the comparison seriously. Under the planner picture the machine improvised. Under the water picture nothing remarkable happened: a flow met an obstacle, pressed on the surfaces available to it, and passed through the opening that existed. Finding what is possible is the only thing a flow does. **FIG. 1 · ONE FLOW, TWO CONTAINERS** *The same particles pour into both containers at the same rate. Only the walls differ, and the walls decide the shape the flow takes.* ## The mechanical basis The comparison has a mechanical basis, which is what separates it from poetry. A language model is a machine for divergence and convergence: it can generate many candidate continuations of a situation, and it can judge candidates against a criterion and keep the ones that pass. An agent runs the pair in a loop. It proposes a step, takes it, looks at what happened, and proposes again, and that loop is a search. A person running the same search tires after a few attempts, so human search is narrow and gives up early. A machine retries for the price of compute, so its search is wide and patient and does not tire, and wide, patient, indifferent probing is how water works on rock. The comparison also fails in one place, and the failure is the most useful part of it. Water converges on its own, because gravity assigns every position a value and the flow follows the values downhill. An agent’s search has no gravity built in. Left alone, the loop wanders, producing plausible steps indefinitely with nothing to tell it which direction is down. Convergence has to be supplied from outside, by something that assigns value to positions the way gravity does: - **A test.** The step just taken passes or fails. - **A rule.** The draft violates it or it does not. - **A payment.** It clears or it does not. - **A person.** They accept the work or send it back. Whatever supplies that gradient is part of the container, which is why the container decides where the flow ends up. **FIG. 2 · SEARCH WITH AND WITHOUT A GRADIENT** *Both panels run the same wandering search. The right one adds a downhill and a pass-fail slot, and the same wandering converges.* ## The container If intelligence is a flow, the important engineering question changes. It stops being how strong the flow is and becomes what shape the flow passes through. Give a model a bare loop, meaning a prompt in, text out, repeat until the model declares itself finished, and it wanders, stops early, and produces work someone has to check by hand. Give the same model tools that act on real systems, written rules that encode judgment, checks that automatically fail work violating the rules, and a record of past work it can read before starting, and the same flow produces work of a different class. We have measured this directly, in a blind comparison where a mid-tier model at a fraction of the price matched a frontier model through identical tooling, an experiment [What compounds in an agent system](https://idylliclabs.com/writing/what-compounds-in-an-agent-system) reports in full. | FIG. 3 · ONE MODEL, TWO CONTAINERS | | | --- | --- | | BARE LOOP | A prompt in, text out, repeat until the model declares itself finished. Wanders, stops early, needs a person to catch mistakes. | | ENGINEERED CONTAINER | Tools that act on real systems, written rules, automatic checks, a readable record of past work. | *The flow is the same in both rows. Everything that differs is owned, accumulates, and survives a model swap.* A container, then, is everything around the model that gives the flow its shape, and it does four things: - **It contains.** It bounds what the agent may touch, spend, and send. Those bounds are the walls the flow presses against. - **It routes.** It decides which work reaches which model, and in what order. - **It filters.** Its checks stop bad work from passing downstream, the way a screen stops debris. - **It amplifies.** A rule written once applies to every run that follows, so a single correction is multiplied across all future flow. No part of this is intelligent on its own. A test suite is not intelligent, a permission boundary is not intelligent, a log is not intelligent. The intelligence of the whole system belongs jointly to the force and to the structure that channels it, and attributing all of it to the force is the mistake the planner picture invites. > A domain-specific language is a small programming language built for one narrow job. A craft rule encoded in one cannot be skipped, because work that violates the rule cannot even be written. The last two functions compound. Each layer of encoded judgment, from a style guide a run must read, to a grammar its output must parse against, to a domain-specific language, removes some of the ways the work could have gone wrong, until the remaining paths are the ones the craft permits. **FIG. 4 · WHAT ENCODED RULES DO TO THE SEARCH** *Twenty-four ways to do the same job enter from the left. Each layer of encoded judgment ends some of them at its gate, and what leaves on the right is only work the rules can accept.* ## People The description fits people too, which is part of why we trust it. A person’s thinking runs through structures the person built earlier: a vocabulary, a set of habits, procedures, tools, the notes they keep. Two people of similar raw ability produce work of very different quality when one of them has spent years building the structures their thinking flows through. A skill unused for a decade is not destroyed, it is a channel the flow stopped visiting, still there and still routable. When we want to think better, the lever that works is the same one that works for machines: change the container. Add a checklist, a sharper tool, a written record, a rule that catches a known mistake. This is why we spend most of our time building containers. The flow arrives from outside: model vendors sell it, its price falls every year, and each release is stronger than the last, on a schedule nobody here controls. The containers are ours. Rules, checks, tools, records, and boundaries accumulate, survive every model swap, and improve with every correction, and when a stronger model ships, a good container turns the added pressure into better work instead of a bigger flood. Designing these structures, deciding what to contain, where to route, what to filter, and what to amplify, is what we think building with AI actually is, and the comparison we ran says most of the quality was already ours to build. REFERENCES - [1] Lex Fridman Podcast #491, [“OpenClaw: The Viral AI Agent that Broke the Internet”](https://lexfridman.com/peter-steinberger/), with Peter Steinberger, 2026. --- # Search that does not tire URL: https://idylliclabs.com/writing/search-that-does-not-tire Area: Foundations · Published: 2026-07-17 An agent loop is a search algorithm: propose an action, carry it out, read the result, adjust. Retrying is the part of search that exhausts a person, and it is the part a machine gets at the price of a model call. An agent working through a real task fails in the open. It tries a command, the command errors, it reads the error and tries a variation, the variation errors too, and a third and a fourth approach follow. To someone watching, this looks like a tool malfunctioning. The reading comes from experience with people: a person’s attempts are expensive, so four failures in a row from a person signal a problem beyond them or a worker who is stuck, and either one calls for stepping in. The machine is not running on the person’s cost structure. Its fortieth attempt costs exactly what its first one did, and it arrives carrying what the thirty-nine failures before it revealed. A person’s fortieth attempt costs whatever remains of their morale. Under the borrowed reading, the agent looks broken. Under its own economics, it is doing precisely what its kind of process does, because the loop an agent runs is a search, and dead ends are the normal product of search. ## The loop is search The loop is simple to state. An AI agent is a model running in a loop: the model proposes an action, a program carries the action out, the model reads the result, and then it decides what to do next. The loop continues until the task is done or the agent concludes it cannot be done. This behavior has a name in computer science. A process that proposes candidate solutions, tests them, and uses the result of each test to choose the next candidate is a search algorithm. The agent loop is a search algorithm running over actions in the world. Each attempt is a probe into the space of things that might work, each failure rules out a region of that space, and each success opens the next step of the path. A searcher that never hits a dead end was handed the answer in advance. ## What a retry costs The steps of this loop cost a person and a machine very different amounts. For a person, proposing an idea and trying it are not the expensive steps. The expensive step is returning to the problem after a failure. Every failed attempt costs attention and morale, both of those are finite, and neither is restored by trying again, so a person’s search stops when their determination runs out rather than when the space is exhausted. People treat persistence as a virtue because it is scarce, and the problems worth solving usually require more of it than the first attempt suggests. For a machine, a retry costs one more pass through the loop, which is one more model call. The failure enters the next call as information, and nothing else changes; the machine holds no discouragement between attempts. > This turns determination from a character trait into a purchasable quantity. It is priced in compute, it is spent by the loop, and it is available in whatever amount the problem requires. Put more compute behind the loop and the search continues, hour after hour, pathfinding toward whatever the available tools make possible. | FIG. 1 · WHAT A RETRY COSTS | | | --- | --- | | A PERSON'S RETRY | Costs attention and morale, and neither is restored by trying again. The search stops when determination runs out rather than when the space is exhausted. | | A MACHINE'S RETRY | Costs one more pass through the loop, which is one more model call. The failure enters the next call as information, and nothing else changes. | | THE CONSEQUENCE | Determination becomes a purchasable quantity, priced in compute and available in whatever amount the problem requires. | *Dead ends are a normal product of search. The price of the dead ends is the price of the search, and only one of the two searchers pays it in a finite currency.* A search paid for this way does more than outlast a person on the problems both were given. It follows paths nobody drew in advance, because a route that exists through the available tools will eventually be probed, whether or not anyone planned it, and this is why agents so often produce solutions their operators never specified. [Intelligence is water](https://idylliclabs.com/writing/intelligence-is-water) describes where that unplanned pathfinding comes from and why the container around the search decides what it produces. **FIG. 2 · A SEARCH THAT DOES NOT TIRE** *No particle knows the way through the baffles, and none of them stops trying. The counter is the number that found the outlet anyway, priced at one retry per attempt.* ## Divergence and convergence The loop described so far searches one path at a time, retrying as it goes. Language models also support a second, wider form of search, and the two forms together account for most of what makes an agent system improve. A language model can generate many candidate solutions to a problem for very little more than the cost of generating one: ten drafts of the same page, eight designs for the same component, twenty subject lines for the same email. A judge, which can be a person, a written check, or another model call, then compares the candidates and selects. We call the first half divergence and the second half convergence: generation spreads the search out across the space, and selection collapses it back to a point. Run once, this produces a better result than a single attempt would, for the ordinary reason that the best of twenty candidates beats the average of one. The step that separates systems that improve from systems that do not is what happens to the selection afterward. When the reason a candidate won is written down as a rule, and every future run is required to read the rule before it starts, the next search begins from a smaller space, because candidates that would have violated the rule are never generated at all. Each round of generating and selecting shrinks the space the following round has to cover. A system built this way gets narrower and more precise with use. A single-attempt system, meaning one prompt, one output, and no record of what was chosen or why, searches the same space at the same width forever, because nothing carries over from one attempt to the next. Both systems can run on the same model. What separates them is whether the selections accumulate. **FIG. 3 · THE NARROWING SEARCH** *Twenty-eight candidate paths enter round one. Each round's selection is written down, so the next round never generates the paths the record has excluded, and the band narrows with use.* | FIG. 4 · SELECTIONS THAT ACCUMULATE | | | --- | --- | | ROUND 1 | A wide fan of candidates. One is selected, and the reason it won is written down as a rule. | | ROUND 2 | Candidates that would violate the first rule are never generated. The fan is narrower, centered where the last selection landed. | | ROUND 3 | Narrower still. Each round searches only the space the recorded rules have not yet excluded. | | SINGLE ATTEMPT | One prompt, one output, nothing recorded. The same wide fan every round, at the same width, forever. | *A system that encodes each selection as a rule searches a smaller space every round. A system that keeps nothing searches the same space forever.* ## The outside judge The selection step needs something real to select against. If the only judge available to the loop is the model itself, the search drifts. A model grading its own output against its own taste makes the output more internally consistent with each round rather than more correct, and the search converges on what the model already believes instead of on what works. Ten rounds of self-review produce an elaborate answer; they do not produce a tested one. The correction has to come from outside the loop: - **A compiler** rejects the code. - **A test** fails. - **A customer** does not reply. - **A measurement** comes back lower than the prediction said it would. Each of these contacts with the world eliminates candidates that no amount of internal reasoning would have eliminated, because the world holds information the model does not. A loop with reality contact converges on solutions that survive it. A loop without reality contact can run forever and converge on nothing, and the compute it burns buys elaboration rather than progress. Agent systems are sometimes dismissed on the ground that they produce careless work, and some do. A system wired to real feedback stops, because every careless output that fails in the world becomes a rule the next run must obey. We build our own systems this way and run our own projects on them, and the same shape repeats in every one. The system generates more candidates than a person would have the patience to write, selects against written checks and recorded judgments, writes the reasons down where the next run must read them, and puts the result somewhere the world can grade it. The retries, the tenth attempt after nine failures, the returning to the problem each morning, all the parts of search that used to be paid for in a person’s determination, run on compute instead. What such a system asks for is the one input that can simply be bought more of. --- # Shipping is merging URL: https://idylliclabs.com/writing/shipping-is-merging Area: Foundations · Published: 2026-07-17 Agents made it nearly free to open parallel branches of work. A branch becomes progress only when its result is merged back into the one thing being built, and the merge is the step that did not get cheaper. One of the businesses we run has a render pipeline that produces a crafted deliverable, and at one point we pointed agents at it in parallel. The agents generated dozens of branches of the same deliverable, each branch a complete working copy that explored a different layout, a different animation, a different structural idea. Every branch contained something worth keeping, and no branch was complete. For a while the pile of branches felt like progress, because the pile kept growing. By the only measure that mattered it was standing still, because nothing had been merged, and the deliverable itself was unchanged. The work got done when the merging began. A single reviewer went through the branches in sequence, took from each the piece that survived judgment, composed the kept pieces into one deliverable, and discarded the rest knowingly. The parallel phase produced the raw material in hours. The finished deliverable came out of the merge, and only out of the merge. ## Splitting is copying The usual belief about parallel work is that it helps only when a task decomposes into independent pieces, so the way to use many agents is to find work that splits cleanly and give each agent a piece. We build tools for directing many agents at once, we run our own projects on those tools, and running them corrected the belief. Almost any work can be split, because splitting is copying: git hands every contributor a complete copy of a codebase that cannot be edited in parallel in place, and the contributions come back as merges. What decides whether parallel work counts is at the other end. A merge is the step that folds one branch of work back into the trunk, meaning the single current version of the thing being built, and the question that predicts whether a group of parallel agents will accomplish anything is whether a merge path exists from each branch back to the trunk, not whether the task looked parallelizable at the start. The two halves of this arrangement are now priced very differently. Opening a branch costs a model call: every prompt starts a new line of work, an agent will carry the line as far as its tools allow, and the number of open branches is limited only by willingness to pay for compute. Merging did not get cheaper, because a merge was never made of labor. To merge a branch someone has to read it, judge which part of it is worth keeping, and fold that part into the trunk without breaking what is already there, and each of those is judgment exercised in sequence. When one half of an operation becomes cheap and the other half stays expensive, the queue forms in front of the expensive half. Anyone who runs agents has seen the symptom: branches, drafts, and near-finished artifacts accumulate, and the thing being built does not move. **FIG. 1 · BRANCHES CONVERGING ON A TRUNK** *Every path on the left is a branch, opened for the price of a model call. The gates are the judgments a branch must survive, and the band converging on the center line is the trunk admitting work in sequence.* ## Amdahl’s law The merge that finished the deliverable ran in sequence, and it had to, because each merge changes the trunk that the next branch lands on. It was also where taste entered the system: reading each branch, deciding what was worth keeping, composing the kept pieces into something coherent. Amdahl’s law, from parallel computing, says that the speedup available from adding processors is bounded by the fraction of the work that has to run in sequence, and in agent work the fraction that has to run in sequence is the merge. The bound is easy to work out with real values. Suppose a deliverable takes ten hours of drafting and one hour of merging. One agent takes eleven hours. Ten agents draft in parallel in one hour, and the merge still takes one hour, so the total is two hours. A hundred agents draft in six minutes, and the merge still takes one hour, so the total is sixty-six minutes. The ceiling is a speedup of eleven times, and no number of additional agents raises it, because the merge hour does not shrink when agents are added. **FIG. 2 · THE HOUR THAT DOES NOT SHRINK** *One bar per fleet size, on one axis of hours. The draft segment collapses as agents multiply; the merge segment is the same length in every row, and the fill spends the same time crossing it.* Every gain past that ceiling has to come from the serial side, and the serial side offers two levers: - **The merge can be made faster.** Written rules and automated checks can judge a branch before a person sees it, so the person’s hour is spent only on what the checks cannot decide. - **The branches can be made more mergeable.** An acceptance condition declared before the branches open tells every agent what the trunk will admit, so fewer branches come back unmergeable. | FIG. 3 · WHERE THE SPEEDUP STOPS | | | --- | --- | | 1 AGENT | Ten hours of drafting, then one hour of merging. Eleven hours in total. | | 10 AGENTS | One hour of drafting in parallel, then the same one hour of merging. Two hours. | | 100 AGENTS | Six minutes of drafting, then the same one hour of merging. Sixty-six minutes. | | THE CEILING | Total time floors at the merge. Speedup is bounded at eleven times regardless of agent count; further gains come only from faster merges or more mergeable branches. | *The merge hour is untouched by every row. Past the second row, agents are nearly free and nearly useless; the remaining gains are all on the serial side.* What the serial step cannot be is removed. One trunk means one ordered sequence of admissions into it, and the admission decision is where the owner’s judgment lives, so handing the whole decision to another agent does not speed the system up; it gives away the property that made the trunk worth converging on. Large engineering organizations already measure at the join: Anthropic, among others, counts progress in merged pull requests rather than opened ones. ## What a merge admits A merge is an acceptance decision, and an acceptance decision needs an acceptance condition, meaning a statement of what must be true for a branch to enter the trunk. Merging without one still lowers the branch count, but the trunk it produces is drifting rather than converging, because each admission moves it and nothing decides the direction. The condition cannot be produced by more thinking. The information that separates a good condition from a plausible one lives outside the system, in what failed, what got a reply, what the world pushed back on, which is the same reason the search loop described in [Search that does not tire](https://idylliclabs.com/writing/search-that-does-not-tire) needs a judge outside itself. The order is therefore fixed: contact with the world first, then the condition, then the convergence. There are also two different ways a branch count falls, and from the outside they look identical. In the first, branches are simply closed: the work inside them is discarded, the trunk receives nothing, and the person has practiced abandoning. In the second, each branch is visited, the part that survives judgment is carried into the trunk, and the rest is discarded knowingly. At agent scale the difference is most of the value, because agent-produced branches are usually partially good: a branch that fails as a whole still contains a section, a structure, or a decision that won. An acceptance rule that admits or rejects whole branches therefore throws away most of what the parallel phase paid for. A convergence is measured by what the trunk gained, and a falling branch count over an unchanged trunk means the branches were abandoned rather than harvested. ## The merge queue Convergence also has to be funded, because generation fills whatever room it is given. A loop that can afford one more draft will produce one more draft, since generating is the cheap operation and selecting is the expensive one, and a search that keeps widening after it has stopped resolving anything is paying to postpone the selection. The device that stops it is a resource limit: a deadline, a token budget, a spend cap. A blockchain calls the same device a gas limit, a budget that halts a computation when it is spent, and the useful refinement for agent work is to reserve the last share of the budget for the merge, because a plan that spends its whole budget opening branches has arranged, structurally, never to collapse them. The artifact for managing the work changes with the prices. A task list controls what gets started, and it was the right control artifact when starting work was the expensive act. A merge queue controls what is admitted to the trunk, in what order, against what acceptance condition, and nothing leaves it except by entering the trunk or being discarded knowingly. The popular multi-agent frameworks are built for the other end: their design surface goes to coordination, to which agent acts next and how agents talk to each other, and a first-class concept of merging results back into one thing is mostly absent, which matches how most agent work is still run. | FIG. 4 · A TASK LIST AND A MERGE QUEUE | | | --- | --- | | WHAT IT CONTROLS | Task list: what gets started. Merge queue: what enters the trunk, in what order, against what acceptance condition. | | WHAT IT MEASURES | Task list: starts, and the states of started work. Merge queue: merges per week, and the age of the oldest unmerged branch. | | ITS FAILURE MODE | Task list: everything in progress, nothing shipped. Merge queue: a visible, aging queue; the debt is on the books. | | WHEN IT WAS RIGHT | Task list: when starting work was the expensive act. Merge queue: when starting became free and the join became the scarce operation. | *The two artifacts gate at opposite ends of the work. When starting became free, the control point moved to the join.* A queue also makes the liability visible. An unmerged branch charges interest: the trunk moves underneath it while it waits, so the cost of integrating it grows with its age. We call the accumulated liability merge debt, and the number we watch is the age of the oldest unmerged branch. The same accounting answers the scaling question people actually ask, which is how many agents a single operator can direct. The number is not set by attention during the parallel phase, because directing that phase costs little. It is set by the operator’s merge throughput, and every branch opened beyond that throughput is debt on the day it is opened. An agent system can run searches that are wide, patient, and nearly free, and what the searches return is raw material. The system that turns the material into a built thing is engineered around the join: an acceptance condition declared before the branches open, a budget whose last share is reserved for the merge, and a queue out of which the only exits are the trunk and the knowing discard. > Everything upstream of that queue runs at the speed of the agents, and the system as a whole runs at the speed of the merge. --- # How to work with many agents URL: https://idylliclabs.com/writing/working-with-many-agents Area: Foundations · Published: 2026-07-18 The mental models that scale one person from ten agents to a thousand: hold macrostates, not microstates, and converge everything into one trunk. It is common now to meet a developer running ten coding agents at once, and the striking thing is where the work actually happens: in their head. They track what each agent is doing, hold how the pieces will fit together, and decide constantly, while the agents generate. The arrangement is genuinely productive at ten agents, and it cannot survive at a hundred, because the integration living in the person’s head grows much faster than the agent count. The ceiling has been measured before, in people. In 1933 a management consultant named V. A. Graicunas took up a question practitioners felt but could not explain: why does an executive with four subordinates hesitate to add a fifth? He answered it by counting what a supervisor actually tracks, which is not four things but four people, every pairing among them, and each person’s view of the rest. Four subordinates put 44 relationships on the supervisor. A fifth raises the count to 100: twenty percent more hands, more than double the tangle. That arithmetic is why spans of control settled near five in every organization that studied them, and a person integrating the work of ten agents in their head is running the same arithmetic on the same head. We run our own fleets at the lab, and two mental models carry all of it. The first: hold macrostates, not microstates. The person states what must be true about the outcome, and the system owns every route to it. The second: however far the work diverges, it converges into one trunk, and human judgment operates at the point of convergence. Neither is a management tip. Both are facts about computation, about what search is and where information comes from, which is why they keep working from ten agents to a thousand. ## Macrostates The first model replaces the question “what should the agents do” with “what must be true when they are done.” The difference shows in two briefs for the same task, a slow endpoint. One brief reads like a plan: open src/api/products.ts, wrap the database call in lru-cache, invalidate on the write path, follow the pattern in search.ts. The other reads like a contract: the tests pass, /products answers a warm request in under 100 milliseconds, and the diff adds no new dependency. The plan is the natural brief for someone who used to do the work themselves, and it fails in two quiet ways. It makes decisions before the search starts: lru-cache is settled in advance, even though the contract version of the same task forbids new dependencies entirely. And its author is the only one who can verify it, since only the author knows whether the route is being followed. Verification that requires the author’s presence is watching under another name. The contract has neither failure. Anyone, human or script, can check its three lines without having seen any of the work. Statistical mechanics supplies the right names for this distinction. A microstate is one exact configuration of a system; a macrostate is the set of configurations that share the properties that matter. Thermodynamics operates entirely on macrostates, measuring the gas and never the molecules, and loses nothing by it. An acceptance condition is a macrostate: it pins properties of the outcome and says nothing about which configuration delivers them. The silence is what scales. Twenty acceptance conditions fit in a head that could never hold twenty routes, and the routes were never the point. Conditions fail at two extremes. Too tight, and the condition is a route again: a person editing an agent’s output word by word has delegated nothing. Too loose, and it constrains nothing: “make it good” admits every output under some reading, and a search under a condition that excludes nothing is a random walk. Conditions that work pin properties and ignore configurations: - The build passes. - Every claim carries a source. - The page loads in under a second. - No written rule is violated. A stranger could check each of those, which is the working test: any condition whose verification requires having watched the run is a route wearing a condition’s clothes. ## Amortized mistakes Micromanagement, in this vocabulary, is attention spent on microstates, and the most tempting microstate is the visible mistake. An agent searches src/ for a file that lives in packages/. The error is obvious to anyone who knows the codebase, the correction costs a sentence, and the detour costs the agent about two minutes. At one agent the correction looks free. At twenty, the two costs sit in different economies: the detour costs compute, which multiplies with the fleet, while the correction costs attention, which does not. Interruption research prices the return from a context switch at over twenty minutes, ten times the detour it prevents. A person who corrects every visible deviation has rebuilt the five-agent ceiling inside the larger fleet. Non-intervention works because a wrong route usually announces itself to the agent that took it. The search returns empty, the test fails, the fetched page lacks what was expected, and the failure enters the next attempt as information. [Search that does not tire](https://idylliclabs.com/writing/search-that-does-not-tire) describes this loop as a search that retries at the price of a model call. Most deviations resolve within minutes, so their cost amortizes across the fleet, while every interruption bills the one resource that has no fleet. **FIG. 1 · INTERRUPT VERSUS AMORTIZE** *Both panels run the same six agents through the same deviations. The top panel routes every deviation through one person's attention, and the runs spend the loop stalled in a queue. The bottom panel leaves each deviation to self-correct about two ticks later, and every run finishes.* Two kinds of mistake escape the amortization argument. The repeated mistake escapes because model calls are stateless: an agent repeating yesterday’s error is reporting that the system never fed the correction back, and the durable fix is a macrostate-level fix, in a form every future run inherits: - **A written rule.** The correction, generalized into an instruction that every future run reads before it starts. - **An automated check.** Where the rule is machine-checkable, an audit that fails the work on violation with nobody present. - **A boundary.** A limit on what a run may touch, spend, or send, so the mistake becomes impossible to express rather than merely discouraged. [Encoding judgment as rules](https://idylliclabs.com/writing/encoding-judgment-as-rules) traces that pipeline, and its economics reduce to one line: watching bills on every recurrence, encoding bills once. The structural mistake escapes for the opposite reason, because it lives above the runs: a wrong decomposition or a wrong acceptance condition corrupts everything downstream, and downstream cannot repair it. Mistakes inside the work amortize. Mistakes in the arrangement of the work compound, and those get corrected the moment they appear. ## Explored space The second model begins with why divergence pays at all, since fanning four agents across one problem looks wasteful next to writing one better prompt. The computational answer: a model samples each token from a distribution conditioned on its context, and a prompt can only shift that distribution. It cannot add information that does not exist. “Weigh the alternatives carefully” cannot supply what happened when the alternatives ran, because nothing has run. Evidence of that kind comes into existence only when something goes and gets it. Four agents sent into four regions of a problem return with what the regions contained: the approach that failed with its error text, the source with its actual claims, the measurement with its value. A synthesis over those reports samples from a distribution moved by evidence no single run could have held, which is the same reason four people who each checked a different store beat one person reasoning about inventory from home. Our operating log states the bet: “20 agents running in different directions collecting data about the world is going to be far superior to one human.” **FIG. 2 · COMPUTATION SPREAD OVER A SEARCH SPACE** *One question fans out into twelve branches; nine are pruned where they fail and three reach the synthesis node. The specks drifting toward the node come from the whole explored region, because a pruned branch reports what a region does not contain, and the synthesis is conditioned on all of it.* > The familiar wrong answer, that the two remaining doors are a fresh fifty-fifty, comes from ignoring that the host’s choice was constrained. The information is in the constraint. Monty Hall is the textbook picture of conditioning on revealed information. A contestant picks one of three doors; the host, who knows the prize and never opens its door, opens a losing one, and switching now wins two times in three, though the first pick never got worse. A constrained process revealed information, and the winning move conditions on it. A synthesis agent reading exploration reports is the contestant after the door opens. The analogy breaks in one place, and the break defines a discipline. The host cannot lie; an agent can, without meaning to. A confident report from a region never visited moves the synthesis distribution exactly as an honest one would, and makes the answer sharper without making it truer. Reports therefore carry receipts: - **Tool output over recollection.** A claim about a system arrives with the command and its output attached, not the run’s memory of having tried it. - **Sources over summaries.** A claim about the world arrives with the fetched document, so the synthesis conditions on what the source says. - **Tests over assurances.** A claim that something works arrives with the passing test. “It works” is a prediction; a test result is an observation. Receipts narrow the room for invented evidence without closing it; a mechanism that closes it has not appeared yet. The shape itself is proven well beyond agents. MapReduce spread computation across machines and combined it in a reduce step, and best-of-N sampling spreads generation across attempts and keeps the winner. In every version, the combining step is worth exactly what the spreading step touched. ## Convergence Divergence, however wide, ends in one place. That is the second model in full: value realizes only when branches fold back into a single trunk, the one current version of the thing under construction, and the fold is where human judgment belongs. Git made the shape familiar. A coding agent checks out a branch, works alone, and returns a pull request, and nothing it did counts until the merge. [Shipping is merging](https://idylliclabs.com/writing/shipping-is-merging) works out the economics: branches became nearly free, merging did not, so the queue forms at the merge, and the person belongs at the front of it, because the merge is the one place where taste compounds. One of the businesses the lab runs delivers crafted video, and its generation fleet produces dozens of incomplete branches, each trying a different structure. The human work there begins at convergence: selecting the pieces that survive judgment, handing the survivors to an agent that stitches them into a single composition, then a denoising loop that regenerates whatever reads wrong until nothing does. No human writes a frame. The finished piece consists entirely of selections. **FIG. 3 · DIVERGE, REVIEW, MERGE** *Eight branches diverge from one trunk. The review gate ends three of them, the five that pass merge back one at a time, and each numbered merge is progress the trunk keeps. The trunk is the only line that ships.* The merge runs one branch at a time, because each admission changes the trunk, and a judgment made against the current trunk is the only kind that stays correct. Mechanical checks run before human eyes, so what reaches judgment is exactly what no rule can decide yet, which is taste. And nothing merges unreviewed, because a pile of unjudged branches is accumulation that feels like progress: bulk-merged, it pushes unjudged work into the trunk, and every future branch inherits it. A fleet that generates ten candidates and promotes winners under a stated condition improves every round. A fleet that keeps everything has produced reading. ## Takeaways The two models, with what falls out of each: - **Macrostates, not microstates.** The person states what must be true in a form a stranger can check, and the system owns every route. - **Mistakes amortize.** Deviations self-correct at the price of compute, which multiplies; only repetition and structural errors escalate, and repetition escalates into rules, not into watching. - **Explored space conditions the answer.** A prompt shifts a distribution; exploration moves it with evidence that did not exist before, and receipts keep the evidence honest. - **Everything converges to one trunk.** Branches multiply freely, checks judge them first, and human taste admits them one at a time at the merge. The two models place the person in exactly two moments: before the work, stating what must be true, and after it, judging what returns. Both moments scale with the number of deliverables rather than the number of agents, which is the entire scaling law. The day keeps its shape while the fleet grows an order of magnitude. One asymmetry compounds in the background. Every repeated mistake becomes a rule, every rule outlives the run that taught it, and each pass starts where the last one ended. The fleet nobody watches gets harder to derail the longer it runs. The whole method fits in a sentence: the person states what must be true, the fleet diverges to explore, small mistakes amortize where they happen, and judgment arrives at the merge, so attention stays free at any fleet size. REFERENCES - [1] V. A. Graicunas, “Relationship in Organization.” Bulletin of the International Management Institute, Geneva, 1933. Reprinted in Gulick and Urwick (eds.), [*Papers on the Science of Administration*](https://archive.org/details/papersonscienceo00guli), Institute of Public Administration, 1937, pp. 183–187. - [2] Gloria Mark, Daniela Gudith, and Ulrich Klocke, [“The Cost of Interrupted Work: More Speed and Stress.”](https://www.ics.uci.edu/~gmark/chi08-mark.pdf) CHI 2008. --- # The agent body URL: https://idylliclabs.com/writing/the-agent-body Area: Foundations · Published: 2026-07-17 Agent products all run on the same models, tools, and prompts, so the parts cannot explain why some improve and hold users while others stay demos. What separates them is the body: what persists between calls, where the boundary sits, and why the boundary is worth having. Every agent product is built from the same short list of parts: a model, a prompt, a set of tools, a few skills, some memory. The parts are rented from a handful of providers, and every builder rents the same ones. Described by their parts, the products are indistinguishable, and most of the time they feel that way too. Yet some of them improve week over week and hold onto their users, while others with an identical parts list stay demos nobody opens twice. The parts cannot be the difference, because the parts are shared. Positioning is the usual explanation, and positioning is not it either. The difference is the body, and the body is the part almost nobody describes. ## What a body is Physically, a body is whatever persists when the model is not running. A model call is stateless. Each time the agent acts, its mind is assembled from scratch, the model and the prompt and the tools it is told about and the documents pulled in on its behalf, and when the call returns, that assembly is gone. Nothing the agent learned during the call survives inside the model, because the model is the same weights afterward that it was before and will be for the next call. So whatever the agent carries from one moment to the next has to live somewhere that outlasts the call. That somewhere is the body. A body is the place where an agent’s history, its learnings, and its standing can exist at all, because a mind that vanishes at the end of every call cannot carry anything. **FIG. 1 · THE MIND SWAPS, THE RECORD HOLDS** *The mind label changes as models are swapped in and out per call. The record underneath never resets, because it belongs to the body, and it is what the next mind finds when it arrives.* ## Why a boundary A place to keep things is necessary but not sufficient. A pile of records with no edge around it is not a body, because nothing about it says which records belong to this agent and which belong to the world. The missing part is a boundary, a line that marks this much as inside and the rest as beyond. The boundary is not a limit grudgingly accepted in exchange for safety. It is the part that does the work, and it works in three ways. The first is that a boundary makes failure mean something. An agent that is for everything has no particular outcome that counts as failing, so no measurement can say whether this week’s version beats last week’s, and there is nothing specific to optimize against. A bounded agent has a defined job, and a defined job produces a defined error. This is the mechanical reason bounded products improve across versions while unbounded ones stay demos: only the bounded one throws off a signal it can climb. The second is that a boundary is what a person can hold in mind. Someone delegates to a tool only when they carry a model of it, and the model they carry is mostly a boundary, this is what it is for and this is where its job ends and mine begins. An agent that offers to do anything asks its user to hold “anything” in mind, and attention slides off anything. Plenty of capable agents go unused for exactly this reason. The capability is real, but nobody can keep a shapeless capability in mind long enough to build a habit around it. The third is that a boundary is what lets anything accumulate. Learnings, records, and standing pile up only because there is an inside for them to pile up in. With no edge, every record the agent produces is more undifferentiated world, indistinguishable from everything else that happened. The membrane is what turns a stream of events into the history of one particular thing. Biology settled all of this early. A cell is not defined by its contents, which turn over constantly, but by the membrane that holds those contents together and controls what crosses in each direction. Remove the membrane and the cell does not become a bigger, freer cell. It disperses into the medium and dies, because death, in thermodynamic terms, is equilibrium with the surroundings: unlimited free exchange, nothing maintained. Theoretical biology gave this a name, autopoiesis, meaning that a living thing is the process that produces and maintains its own boundary.[1] The free-energy principle in neuroscience makes the matching claim from the other side: a system persists exactly to the degree that it keeps its inside statistically different from its outside.[2] An agent that can touch anything and be touched by anything is a cell without a membrane, in free exchange with everything and maintaining nothing. This is also why bounded things specialize, and specialization is where power comes from. A liver is enormously capable inside its boundary and useless outside it, and the boundary is precisely what makes it improvable, because a defined job produces a defined error. An agent product obeys the same rule. The ones that get better are the ones that maintain a membrane. **FIG. 2 · A BOUNDARY AND ITS GATES** *The flow inside presses on the whole boundary all the time. The boundary holds everywhere except at its two gates, so the only flow that leaves is flow that leaves through a permitted opening.* ## Why unity A boundary settles what is inside, and a body usually has several things inside it: a part that holds money, a part that holds correspondence, a part that holds the work queue. For these to be one body rather than a heap of parts, they have to present themselves to the world as a single continuing thing, with one name and one history. Unity is not a nicety here, it is what makes the accumulation worth anything. Standing attaches to an identity, so if the parts answer to no single name, there is no one for a counterparty to come to trust. A history is the history of something, so if the parts keep separate records, no single thread gets longer and more valuable with use. Learnings improve one improvable thing only when there is one thing for them to improve. The parts hold together as a unit by coordinating through shared state: an action taken in one part is written where the others read it, so what one part did becomes something all of them know. That shared record is how many organs behave as a single agent, and it is why the agent can answer for what it did last week, since the week is in the record and the record is the one the whole body shares. ## Mind and body With the boundary and the unity in place, the mind and the body separate cleanly. An agent, seen whole, has a mind and a body. The mind is the model plus everything assembled into its context for a single call: the prompt, the tools it is told about, the skills and instructions it is handed, the documents retrieved on its behalf. The mind is stateless and temporary, put together for one call and gone when the call returns. The body is everything that persists between calls, bounded by the membrane and unified under one name. A body answers five questions: - **How does the agent persist?** What state, records, and accounts remain when no model is running? - **What organs does it have?** Which functional parts own which piece of its domain? - **How do the organs coordinate?** Through what shared state does one part find out what another did? - **What functions do the organs perform together?** Which capabilities exist only in the composition, the way circulation exists in no single organ? - **How does the agent maintain unity?** What makes its parts present themselves to the world as one continuing thing with one name and one history? ## The parts In practice the answers recur, so a body has a recognizable parts list. At the center is a durable record, and the relationship between agent and record runs opposite to intuition: the agent does not have a memory so much as the memory has an agent, one record that successive minds visit, read, and extend. Around it sit the organs, each owning one piece of the domain and coordinating through that record rather than through each other. There are inbound channels, an inbox and webhooks and triggers, through which events reach the body whether or not any model is running, because an agent that can only be prompted has half a body at most. There are its own addresses and accounts, the surfaces through which it acts and is known, because standing with other parties attaches to an address rather than to a model. And there are constraints: spending limits, sending rules, approval gates in front of anything irreversible. The body is defined by these as much as by its capabilities. A wallet with no spending rule is an open channel through which one bad call becomes an unbounded loss, and the rule has to arrive with the wallet, because adding the constraint after the first incident means the first incident was the design process. Constraints are also how the boundary stays a boundary: a limit the body enforces on itself is the membrane doing its job from the inside. Finally, there are written rules the body can be reprogrammed with, which is where human review pays off, since a rule written down today is a part of every mind assembled tomorrow. | FIG. 3 · AGENT — COMPOSITION | | | --- | --- | | THE MIND | The model plus the context assembled for one call: prompt, tools, skills, retrieved documents. Expires when the call returns. | | DURABLE RECORD | One record of everything that has happened, which successive minds visit, read, and extend. Persistence. | | ORGANS | Functional parts that each own one piece of the domain: money, correspondence, the work queue. Division of the domain. | | SHARED STATE | The record through which one organ finds out what another did. Coordination. | | INBOX & WEBHOOKS | Channels that receive events whether or not any model is running. Inbound events. | | ADDRESSES & ACCOUNTS | The surfaces through which the agent acts and is known. Presence and standing. | | LIMITS & GATES | Spending limits, sending rules, approval gates on anything irreversible. The boundary, self-enforced. | | WRITTEN RULES | Rules written today that are part of every mind assembled tomorrow. Reprogrammability. | *The mind is assembled from scratch for every call, and the body is what the next mind finds when it arrives. Skills and prompts live in the mind while a call runs; their durable copies are written rules stored in the body.* ## Four products Real agent products sort cleanly under this frame. A chat assistant has the largest mind and the thinnest body. The mind is a frontier model with conversation context and a few injected memories; the body is a conversation log and a user profile, with no addresses, no accounts, and no organs. The thin boundary explains both of the product’s famous properties. Everyone knows what a chat is for, so the product binds attention instantly, and almost nothing accumulates inside the membrane, so users defect the day a better model ships. A coding agent has an ordinary mind (a model, a system prompt, file and shell tools) and a superb borrowed body: the repository, with its working tree, branches, tests, and continuous integration. The repository explains why coding agents worked before most other agents did. A repo is a ready-made membrane with a built-in error signal, because a failing test is exactly the bounded, optimizable failure that an unbounded agent never gets to have. An autonomous business agent inverts the chat assistant. Its mind is swappable and unremarkable, and its body is the heaviest in the list: a payment account, an inbox, a domain, a ledger of records, and the credentials that hold them, with approval gates in front of spending and sending. The boundary is literally the credential set, since inside is whatever the body holds a key to. That is why this class can run unattended, because events arrive through its own channels, and why its progress survives a change of model, because everything accumulated lives on its side of the membrane. An embedded product agent, the assistant that lives inside an existing application, has a purpose-built mind and no body of its own. It borrows the host’s: the host’s database, permissions, queues, and users. The borrowed boundary explains the whole trade. The agent is good on its first day, because it inherits a mature body, and it can never leave, because the body was never its own and the identity the world sees belongs to the host. | FIG. 4 · FOUR KINDS OF AGENT PRODUCT | | | --- | --- | | CHAT ASSISTANT | The largest mind, the thinnest body: a conversation log and a user profile. Binds attention instantly, and users defect on the next model release. | | CODING AGENT | An ordinary mind and a superb borrowed body, the repository. A ready-made membrane with a built-in error signal, the failing test. | | BUSINESS AGENT | A swappable mind and the heaviest body: payment account, inbox, domain, ledger, credentials, approval gates. Runs unattended, and its progress survives a model swap. | | EMBEDDED AGENT | A purpose-built mind and a borrowed body, the host’s. Good on day one, and it can never leave, because the identity belongs to the host. | *A thin or borrowed body is a design position, not a defect. Each row records consequences, not rankings.* ## Three questions Any product in this list, including one you are building, can be checked with three questions: - **Does the agent know what it did last week?** A body holds its own history; a visiting mind holds nothing. - **Does an action taken through one surface show up everywhere?** Every other place the agent lives should see it, because the organs coordinate through shared state. - **Does it defend its own invariants?** Refusing, on its own, the action that would cross one of its limits. Three yeses is a body. Fewer, and what exists is a mind making visits: capable, stateless, and gone when the call returns. The minds are converging anyway. Every builder rents them from the same few providers at the same prices, and a better one is available to every competitor on the day it is available to you. The body is the part that cannot be rented, which is why the body is what a company actually makes. The direction this points is many such bodies existing digitally, each maintaining its own boundary and finding its own market. REFERENCES - [1] Humberto Maturana and Francisco Varela, *Autopoiesis and Cognition: The Realization of the Living*. D. Reidel, 1980. - [2] Karl Friston, [“The free-energy principle: a unified brain theory?”](https://www.nature.com/articles/nrn2787) Nature Reviews Neuroscience, 2010. --- # Nothing grown in isolation survives transplant URL: https://idylliclabs.com/writing/nothing-grown-in-isolation-survives-transplant Area: Foundations · Published: 2026-07-17 Shipping speed and growth location look like one variable and are two: a team can launch every week while everything it builds grows against its own model of the market. Readiness is a product of contact with the destination, not a precondition for it. A seedling started indoors and moved outside too fast dies within days, and not from the cold. It dies because everything about it, the thin skin of its leaves, the soft stems, the roots that never met wind or a dry afternoon, was built for the room it grew up in and not the ground it was moved into. Gardeners have a name for the fix. You harden the seedling off: you carry it outside for a little longer each day while it is still growing, so that by the time it goes into the soil it was already, in every way that matters, growing there. The same thing is true of the things we build. ## Two variables The standard advice against building in isolation is to launch fast. The advice answers the wrong variable, because shipping speed and growth location look like one variable and are two. Shipping speed measures how quickly artifacts leave the building. Growth location is where the organism’s roots are while it is being formed, and the two vary independently. A team can ship every week and still be growing in isolation, because rapid iteration against its own model of the user is a greenhouse running at a high frame rate. A team can ship nothing for a year and be growing in full contact, because one real customer conversation a week, sustained over months, is contact. The variable that decides the transplant is whether signal from the destination flows in and out while the thing grows, and shipping speed does not measure that. The two variables come apart cleanly in a pair of founders. The first builds alone for a year and ships a polished product: fast-looking output, and no outside signal during any month of its formation. The second ships nothing for that same year and has fifty conversations with real customers: slow-looking output, and signal in and out the whole time. On launch day the first product is transplanted for the first time, into ground it has never touched. The second has been growing in its destination since the first conversation. Speed did not decide the outcome. Where the growth happened decided it. **FIG. 1 · TWO PLACES TO GROW** *The greenhouse is sealed: its flow circulates against its own walls, and nothing enters or leaves. The open ground is the same container with openings, and the counter is the outside signal that reaches what grows in the middle.* ## Selection pressure Growth location decides the outcome because anything that grows, grows under a selection pressure, and it becomes fitted to that pressure and to no other. When the pressure comes from the destination environment, the fitness is real. When the pressure is synthetic, meaning a builder’s own judgment, an internal model of the user, or an eval the builder wrote, the fitness is synthetic too, and the thing being grown becomes very good at satisfying a world that does not exist. A founder can spend two years alone making an architecture more and more satisfying to the only judge in the room, and learn less from those two years than two weeks of embarrassing real usage would have taught, because the two years were graded by the wrong judge. The two years do not feel like a mistake while they are happening. With no signal flowing in or out, the internal model of the market is the only judge of the work, and the drift between the model and the territory cannot be seen from inside, which is why two isolated years can feel like steady progress in every one of their months. None of this argues against preparation. Surgeons train before they operate, pilots train on simulators before they fly, and the training works because a flight simulator copies its selection pressure from the destination: real aerodynamics, real failure modes, real instrument behavior. Preparation whose pressure is copied from the destination is training. Preparation whose pressure is disconnected from the destination does something else: it adapts what is being grown, with great discipline, to a world that is not there. ## No private curriculum The disconnection matters most for the capacities that decide survival in a market, because those capacities cannot be prepared privately at all: - **A tolerance** for outreach and rejection. - **A feel** for what a customer will actually pay for. - **The words** customers use for their own problem. - **A calibrated sense** of which replies are interest and which are politeness. None of these can be acquired by thinking harder. Thinking that runs upstream of contact elaborates the model without correcting it, so the model gets more detailed and no more accurate, because nothing outside it pushes back. Contact supplies what private effort cannot, which is the questions you did not think to ask, the surprises, and the verdicts. There is no book and no simulation that transmits them, in the same way that reading about making wheels is not making wheels. Some capacities do have a private curriculum. A person can learn to program alone in a room, and many good programmers did exactly that. This is what makes private preparation tempting: the first skills a builder acquires are usually the ones that private study genuinely can supply, so it is natural to assume the remaining skills work the same way. They do not. The capacities with no private curriculum are the market-facing ones, and the market-facing ones are the ones the transplant is graded on. ## Naming a person The most common way to feel in contact without being in contact is to ask about the destination at the wrong altitude. “What does the market want?” sounds like contact, and it cannot be searched: the subject is unbounded (which humans, in what context?), the predicate is undefined (wanting, in what sense?), and no observation settles it. A search needs constraints to produce a gradient, and a question with no gradient gives the search no direction to move, however much intelligence is applied to it. Naming a person collapses every one of those dimensions at once. “Did Dana, Marcus, and Priya, who sat through the demo last Tuesday, say they would pay forty dollars a month?” has one observable subject, a vocabulary made of whatever words the three of them actually used, a number that will be accepted or refused, and an obvious next step, which is one more conversation. The person has to be real, with a name and an inbox, because a persona answers with whatever the internal model already believed. | FIG. 2 · THE SAME QUESTION AT TWO ALTITUDES | | | --- | --- | | THE SEGMENT FORM | “What does the market want?” The subject is unbounded, the predicate is undefined, and no observation settles it. A space with no gradient, however much intelligence is applied to it. | | THE NAMED FORM | “Did Dana, Marcus, and Priya say they would pay forty dollars a month?” One observable subject, a vocabulary made of words real people used, a number that will be accepted or refused, and an obvious next step. | | THE COLLAPSE | Naming a person collapses subject, vocabulary, price, and next step at once. The search acquires a direction to move. | *The same question, at two abstraction levels, is the difference between a search with a gradient and a search with nowhere to go.* A business is not built for three people, and the objection that a market is bigger than Dana is correct. The named person is not the market; she is the coordinate at which the search becomes runnable. Generalization happens afterward, outward from a convergence point that actually exists, instead of beforehand, downward from an abstraction. Starting from the segment hands back the unsearchable question. Starting from the person produces a foothold, and from the foothold the segment becomes visible for the first time. **FIG. 3 · THE QUESTION AS A SEARCH SPACE** *Both panels run the same wandering search. The segment form is a closed box with no downhill, so nothing ever settles. The named form adds a gradient and a slot that will be accepted or refused, and the same wandering acquires a direction.* ## Where agent bodies grow An AI agent, as we build them, has a mind and a body, a decomposition [The agent body](https://idylliclabs.com/writing/the-agent-body) develops in full. The mind is the model, and it is rented. The body is everything that persists between calls, meaning the accounts, the records, the written rules, and the standing with the parties it deals with. The field mostly grows agent bodies in sandboxes: synthetic benchmarks, eval harnesses, simulated users, replayed logs. A body optimized against a sandbox is fitted to the sandbox’s selection pressure, and a sandbox’s pressure is disconnected from any market’s by construction. It is the two-year founder rebuilt in software, and it dies at transplant for the founder’s exact reason. Sandboxes keep the work that has a private curriculum: whether the code runs, whether a tool call parses, whether the agent stays inside its guardrails. Those verdicts are the same in a sandbox as anywhere else, which is the flight-simulator condition. The capacities that decide whether an agent business survives, such as feel for a market, customer language, and calibrated willingness-to-pay, have no sandbox curriculum, so the businesses our agents run are grown against real markets from the first day: incorporated, holding a bank account, taking real payments, sending real email, absorbing real rejections, with a person approving anything consequential. A business grown this way is small and unimpressive early, the way a seedling hardened off on a windowsill is smaller than one forced in a warm room. It also never has a launch day in the transplant sense, because by the time anyone would call it launched, it has been growing in its destination the whole time. --- # What compounds in an agent system URL: https://idylliclabs.com/writing/what-compounds-in-an-agent-system Area: Persistent state · Published: 2026-07-17 In a blind comparison, a mid-tier model matched a frontier model on every dimension our tooling had encoded. The model is the part of an agent system to plan on replacing, and what compounds is the owned layer: the history, the verdicts, the encoded craft, and the standing it earns. Anyone who sets up AI agents to do real work has to pick a model to run them on, and the price difference between models is large enough that the choice matters. The natural assumption is that the model is the intelligence, the intelligence does the work, and so the system will be about as good as its model, which makes the correct move buying the most capable model available and worrying about cost later. We build tools for directing many agents at once, we run our own projects on those tools, and we held this assumption too: for months we scheduled work around frontier-model quality. A comparison we ran in one of those projects ended it. The briefs were identical and the tooling was identical, with a frontier model producing half of the drafts and a mid-tier model at a fraction of the price producing the other half, and the drafts were graded blind, so the reviewer never knew which model produced which draft. > Blind grading matters because a reviewer who knows which model produced a draft reliably finds the quality they expect to find. The mid-tier model matched the frontier model. On every dimension the tooling had encoded, the drafts were indistinguishable, and drafts from both models passed the project’s automated checks at the same rate. ## Where the quality lived > A domain-specific language is a small programming language built for one narrow job. Ours encodes decisions a craftsman would otherwise make by hand on every draft, so a run cannot skip them. Both models were writing through the same machinery. The project produces a crafted deliverable, and for weeks its quality had been carried by an accumulating layer of tooling rather than by any particular model: a domain-specific language that encodes the craft decisions, written checks that automatically fail a draft when it violates a rule, and a library of graded past examples that every new run reads before it starts. The one place the frontier model still won was a fine-grained interaction sequence the tooling had not yet encoded, and that exception located the craft precisely: the quality lived in the machinery we had written down, and the model only mattered where the machinery stopped. **FIG. 1 · SAME TOOLING, TWO MODELS** *Both panels run the same brief through the same tooling. The counts cascade to the same values, because the quality lives in the machinery the two models share.* ## Rented and owned The result sorts an agent system into two kinds of parts. The model, together with the context assembled for one call, is rented: it is paid for per call, it is supplied by a handful of vendors at prices that keep falling, and it remembers nothing from one call to the next. Everything else persists between calls and is owned: the accounts and credentials the system holds, the append-only history of everything it has ever done, the graded judgments recorded on its past work, the craft encoded as rules and checks, and the standing it has built with the people it serves. | FIG. 2 · WHAT LIVES WHERE | | | --- | --- | | MODEL | The model plus the context assembled for one call. Rented per call, from vendors everyone shares. Remembers nothing. | | EVENT HISTORY | Append-only record of every action the system has taken. Owned. Grows only through operation. | | VERDICTS | Graded judgments on past work, with the reasons attached. Owned. Compiled into rules over time. | | ENCODED SKILLS | Rules, checks, and procedures every future run must read. Owned. Enforced automatically where machine-checkable. | | STANDING | Accounts, domains, payment rails, and reputation with customers. Owned. Slowest to build and hardest to replace. | *The first row is rented per call and replaced without loss. Every row below it survives a model swap.* Models are interchangeable for a structural reason. Everyone can rent the same models at the same prices, so whatever advantage a model confers, it confers on your competitors in the same month. The gap between the best model and the tier below it also keeps narrowing, and it narrows fastest exactly where tooling can carry the difference, which is what the comparison measured. > An advantage that everyone can buy is a cost, not a moat. The owned parts have the opposite property: nobody else can rent your event history, your recorded judgments, or your standing with customers, because those exist only inside your system and accumulate only through its operation. Two questions decide which side a given part of your system falls on: - **If the current model disappeared tomorrow,** would the system still have its identity, its resources, its history, its obligations, and its reputation? - **Could a different model take over the same state** and continue the same work without customers noticing? When both answers are yes, model choice stops being an identity decision and becomes a pricing decision, and every model release becomes good news, because the improvement arrives into a system that keeps everything it has learned. **FIG. 3 · WHAT A CALL KEEPS, WHAT THE SYSTEM KEEPS** *The left panel assembles a context for one call and wipes it when the call returns; its counter never moves. The right panel is everything that persists between calls, and it only counts up.* ## Building for the swap Most of the market is currently building the other way around. Products sell the model’s intelligence, treat the money and the records as plumbing, and store their accumulated judgment in prompts and fine-tunes that die with the model that carried them. If the comparison result generalizes, and our operating experience says it does, that inventory is worth less than it looks, because the part being sold is the part that commoditizes. For anyone building agent systems, the practical consequences are direct: - **Move judgment out of the prompt and into written rules.** A prompt lasts one call and evaporates. A rule file that every future run must read persists, and it compounds. - **Record every verdict where the next run can read it.** A judgment given once and thrown away has to be paid for again tomorrow. Graded examples with the reasons attached are the cheapest asset an agent system can accumulate. - **Keep the event history append-only and owned.** The history is what the next model reads to continue the work, and it is the one record that cannot be reconstructed later. - **Re-run the blind comparison on a schedule.** Tooling keeps absorbing craft, so the cheapest model that passes keeps changing. The comparison is how you notice. When a better model ships, we swap it in, and the system keeps its history, its rules, and its customers. The model is the one part of the system you should plan to replace. --- # The value of the harness URL: https://idylliclabs.com/writing/systems-appreciate-prompts-depreciate Area: Persistent state · Published: 2026-07-17 A model in a chat window gives you drafts you still have to check. The same model inside a real pipeline produces finished work. This essay is about the harness, everything built around the model, and why each new model release multiplies whatever you have built there. A chat window is a harness, and it is the minimal one: a text box, a transcript, and the model, with nothing read before generation, nothing checked after it, and nothing kept between sessions. The same model dropped into a production system, one where every run reads written rules before generating, compares its work against graded examples, passes automated checks, and moves through a staged pipeline, produces work of a different class. In the chat window the model returns a draft that has to be steered, corrected, and checked by hand. In the production system it returns finished work. The intelligence is identical in both cases, so the whole difference in output is the value of the harness, made visible. The gap has a mechanical explanation. Anyone who has worked with a language model has told it to think harder, and the model complies in the only way it can, by producing text that sounds like harder thinking: longer sentences, more qualifications, a more deliberate tone. The instruction buys a style, because a style is the only thing an instruction can buy in a bare harness. No dial inside the model maps “think harder” to more computation, so the words purchase the register of effort rather than the effort. The dial exists one level up. A pipeline can route a hard problem through more passes, more parallel candidates, and stricter checks, which is the extra computation the instruction was reaching for. The bare harness has a second limit: nothing in it persists. A chat is one context window with no memory across sessions, so a correction given there improves the draft on screen, and then the window closes and the correction is gone. Both limits point the same way. A chat honors a request once, as a style, for one reply. A harness enforces the same demand on every run, as a mechanism, indefinitely. ## What a harness is In our systems the harness is concretely a codebase: the files, checks, examples, and pipeline code that exist between model calls and persist while the model inside is swapped per call. An effective harness, in our experience, has four parts, and the parts only work when they are versioned together as one unit. - **A skill file** is the written rules of the craft, and every run reads it before generating anything. - **A golden set** is a library of past outputs that were accepted, stored with the grades and the reasons, and new work is compared against it. - **A verifier** is a set of automated checks that fail a draft when it violates a rule, and it runs without a human present. - **A rejection ledger** is the record of failed outputs with the reason each one failed, and it is where the next rule comes from. Around the four sits the pipeline: staged passes, checks between the stages, and a rule for how many parallel attempts get reduced to one accepted result. The four pieces reference each other, which is why they version as one unit. A golden set without its verifier drifts, because nothing fails the examples that stop deserving their grades. A verifier without a ledger stops growing, because the record of what failed is where new checks come from. A rule without the example that motivated it gets argued with, by humans and by models alike. A change to any piece updates the others. Written checks sound like evals, and an eval suite is one of the four pieces. The difference is where corrections land. Most eval practice grades the model and leaves the grader’s reasoning in a spreadsheet or a chat thread. When the four pieces are versioned together, every correction lands in an object that every future run must read, so nothing the operator learns evaporates. [The agent body](https://idylliclabs.com/writing/the-agent-body) describes an agent as a stateless mind visiting a persistent body, and the harness is the load-bearing part of the body: the part that carries what the operation has learned. Minds are rented from vendors and swapped without ceremony. The harness stays, and whatever the system learns is either written into it or lost. ## What a release does A model release is intelligence arriving from outside into whatever harness exists when it lands. In a bare harness the gain is mostly wasted: the smarter model re-derives context it was never given, repeats failures nobody recorded, and produces variance nobody can review. In a built harness the gain converts: the smarter model reads the same rules, passes the same checks, and clears work the previous model could not, often in a single pass. Work that needed steering under the old model becomes a one-shot process under the new one, and the reason is not the new model alone. The value was pre-baked into the harness while the old model ran, and the upgraded model collects it on arrival. **FIG. 1 · THE HARNESS ACROSS THREE RELEASES** *The vertical lines are model releases. The bright line is the harness's accumulated examples, rules, and checks, and the jump at each line is a new model collecting value banked before it existed. The dim line is tuning tied to one model's specific weaknesses, written off at the same line.* Not everything in the codebase is the value. A smarter model can regenerate scaffolding and glue from scratch, and it will, so the generic code depreciates along with the tuning. What no model can regenerate is the record of your selections against reality: which outputs you accepted, which you rejected and for what reason, which checks encode constraints of your actual domain. That part of the harness is data about judgment rather than code, it accumulates only through operation, and it is the part a release multiplies. ## A natural experiment During a client engagement we had temporary access to a frontier model. Partway through, the access was revoked, and some weeks later it came back. The revocation turned out to be a controlled experiment on where the value of the work lived. While the access was gone, the work continued on weaker models, and the effort went into the harness. The frontier model’s earlier outputs were graded into a golden set. The patterns that made those outputs good were extracted into written rules. The checks that would have caught the weaker models’ failures were built and wired into the pipeline. When the access returned, the same model, no smarter than before, cleared the entire accumulated backlog almost immediately, at roughly thirty percent of the usage it had needed earlier, because it now flowed through the pipeline that had been built in its absence. The model was identical before and after, so the only variable that changed was the harness. Before the revocation, a strong model in a thin harness produced good outputs slowly, each one hand-steered by the operator. After it, the same model in a built harness finished the backlog in one pass, and prompting effort went down rather than up, because the pipeline reads the rules so the operator does not restate them. A returning model is, from the harness’s side, exactly what a release is: intelligence arriving from outside into structure that was ready for it. Every future release replays this experiment with a stronger model in the returning role. The harness could only be built because one real deliverable had already shipped: the graded examples had to come from work that had made contact with a real recipient, because a golden set graded against imagined standards records taste nobody has tested. A blind comparison we ran on the same class of system points the same way, and [What compounds in an agent system](https://idylliclabs.com/writing/what-compounds-in-an-agent-system) reports it in full: the model mattered only where the tooling stopped. ## Compiling judgment The corrections a person gives in chat are the most valuable data they produce and the least preserved, because by default they die with the context window. Telling a model in chat to stop using a phrase fixes the draft in front of you and nothing else. Writing the same instruction into the skill file that every run must read, and storing the rejected draft in the ledger with the reason attached, fixes every draft produced after that day. The correction is identical in both cases. What differs is its half-life, which goes from one call to permanent when the correction lands in the versioned unit instead of the conversation. The same move works one level higher. When outputs are wrong, the instinct is to argue with each one in chat, and the productive response is to change the rule, the example, or the check that allowed the output, which closes the whole class of failure at once. **FIG. 2 · CORRECTIONS LANDING IN THE VERSIONED UNIT** *Each row is one judgment that would have died with a context window, recorded instead where every future run must read it. The counter is compiled judgment, and it only goes up.* Run over months, the loop amounts to compiling your own judgment into an executable form. Each rejection enters the ledger with its reason, each correction becomes a rule, each accepted output joins the golden set, and the compiled judgment then runs in every parallel agent at once, on every future model, with nobody present. > The frontier is rented. What you distill from it is owned. The obvious worry is that compiling your judgment automates you away. What gets compiled is the settled part of the judgment, the rules you no longer want to restate, and compiling them moves the review upward: the person spends their time only on what the harness cannot yet check, and each session at that level produces the next rule. The operator’s position in the loop is permanent even as its altitude rises, and the skill practiced there shifts from producing outputs to producing the harness. | FIG. 3 · WHAT THE HARNESS KEEPS | | | --- | --- | | CLEVER PROMPT | Works once, on one model. Its durable form: the named rule extracted from why it worked. | | FINISHED ARTIFACT | Shipped and superseded. Its durable form: the graded example library it was selected into. | | FRONTIER ACCESS | Rented, and revocable. Its durable form: the rules and checks distilled from the frontier model’s outputs. | | THIS WEEK'S OUTPUTS | Replaced by next week’s. Their durable form: the record of this week’s failures, with reasons. | | REQUESTED REGISTER | Asked for in chat, honored for one draft. Its durable form: the pipeline structure that forces it on every run. | | CHAT CORRECTIONS | Gone when the context window closes. Their durable form: the versioned unit every future run must read. | *Every row's first form dies with the session or the model generation it belonged to. The durable form lives in the harness, survives the model swap, and is multiplied by it.* ## Compute market fit A practical test for a harness is to drive more compute into it and watch what comes out. An effective one converts more compute into more finished, accepted work, a property we call compute market fit. A bare harness fails the test: more runs produce more variance to hand-review rather than more accepted output, because nothing converges the candidates. Verifiers and golden sets are what turn parallel generation from a review burden into a search: candidates that fail the checks are discarded mechanically, candidates that pass are compared against the graded examples, and the human sees only the survivors. The property that absorbs a smarter model is the same property that absorbs more compute from the current one, because both are more intelligence flowing in, and the harness is what turns intelligence flowing in into work coming out. Model releases will keep arriving, on a schedule nobody downstream controls, and each one will land in whatever harness exists at that moment. In a bare harness a release is a slightly better chat. In a built harness it is a payout: the backlog clears, the cost per accepted result drops, and work that needed supervision becomes one-shot. The work that pays is building the harness the next model will flow through. --- # Silence is a signal URL: https://idylliclabs.com/writing/instrumenting-agent-work Area: Measurement · Published: 2026-07-17 Most work an agent sends into the world gets no reply, and a bare no-reply teaches the system nothing. Splitting the outcome into stages, received, opened, clicked, answered, turns silence into an error signal the system can actually learn from. An agent loop is a search: it proposes an action, takes it, reads the result, and proposes again. [Search that does not tire](https://idylliclabs.com/writing/search-that-does-not-tire) describes why the loop needs a judge outside itself, and [Intelligence is water](https://idylliclabs.com/writing/intelligence-is-water) describes what the judge supplies, a gradient that tells the search which direction is down. Inside our own systems we build the judges deliberately: a test passes or fails, a draft violates a written rule or it does not, and a merge goes through when the work meets an acceptance condition set in advance, which [Shipping is merging](https://idylliclabs.com/writing/shipping-is-merging) argues has to come from outside the work. Every one of these is an error signal: a measurement that changes depending on which part of the work failed. The moment work leaves the system, the judges stop. In one of our projects, agents research a recipient, build a piece of work specifically for that recipient, and send it with a short note. The dominant outcome, as with any outreach, is silence: most of what gets sent receives no reply, and this stays true even when the work is good. Silence is not an error signal, because it does not change depending on what failed. A message that never arrived, a message that arrived and was never opened, and a message that was opened, read to the end, and declined all produce the same observation, which is nothing. The usual response is to treat that silence as one fact meaning no, and then the only lever left is volume, which is how agent outreach becomes spam. The real problem sits earlier: a search fed by silence has no gradient to follow, so the loop cannot converge no matter how many attempts it can afford. No better sensor fixes this, because there is nothing out there to sense. What we can change is the shape of the outcome itself: the path from sending to reply decomposes into stages that each either happened or did not, and each stage can carry its own measurement. The decomposition is what creates the error signal. ## Decomposing the outcome > Signed means the token is generated with a secret key, so a stage event cannot be forged or attached to the wrong recipient. A reply is the last event in a chain, and every earlier link is a separate fact about one recipient: the message arrived, they opened it, they clicked the link to the work, they got far enough into the work to judge it, they answered. To measure the chain link by link, every piece of work we send carries its own unique signed token, a link parameter that identifies exactly one recipient. Beacons, small requests that fire when the page loads and when the work starts playing, record the early stages, and we log replies by hand, because they arrive in language and carry sentiment a sensor cannot read. Each transition lands as a typed event against that one recipient, in the same append-only log the rest of the project runs on, so the funnel report is a query rather than a spreadsheet someone maintains. **FIG. 1 · THE MEASURED PATH FROM SENT TO REPLIED** *One wave of two hundred sends, cascading stage by stage. The dots falling away between rows are the recipients lost at that exact stage, which is where the fix belongs.* Measured this way, the same silences separate into different failures, and each failure has its own fix: - **Sent but never opened.** The work was never judged at all. The failure belongs to the first sentence of the note, the channel, the send time, or the sender’s standing on that channel, and improving the work itself would change nothing. - **Opened but never started.** The page carrying the work failed, which is rare and usually mechanical: load weight, a broken embed. - **Started and abandoned early.** The work’s opening failed, and the adjustment is to move the strongest material earlier or make the whole thing shorter. - **Finished with no reply.** The work succeeded and the ask failed. This is the only case where a follow-up is justified, and it goes only to the people who finished, with a single concrete question, because chasing the people who never opened anything on the same channel is how sender standing gets destroyed. | FIG. 2 · WHERE THE FUNNEL LEAKS | | | --- | --- | | SENT → NOT OPENED | The note or channel failed; the work was never judged. Adjust the first sentence, the channel, or the send time. | | OPENED → NOT STARTED | The carrying page failed. Check load weight and the embed; this stage should be near-lossless. | | STARTED → ABANDONED | The work’s opening failed. Move the strongest material earlier, or shorten the whole piece. | | FINISHED → NO REPLY | The ask failed, the work did not. One follow-up, only to finishers, with a single concrete question. | | REPLIED | A reply arrived. Its sentiment is recorded by hand, and a person takes over when the reply asks for one. | *Each row is one measured transition. Read down for the first stage with a disproportionate drop; that is the stage to fix, because fixing a later stage does nothing for the recipients already lost above it.* Nothing about the recipients changed, and the same silences arrive in the same volume. What changed is that each one now lands at a stage, and a silence with a position in the chain is an error signal: it points at the part of the system that failed, and just as usefully at the parts that did not. The decomposition also pulls apart two judgments that a bare silence merges into one. A recipient who finished the work and never replied has judged the work well enough to finish it and has declined the ask, and those are separate readings on separate parts of the system. ## One change per wave > A sampling period, in a control system, is the interval between readings of the thing being controlled. The wave plays that role here. An error signal also has to stay attributable: when a stage rate moves between readings, we need to know which change moved it. Sending continuously blurs every reading, so we send in waves and treat each wave as a sampling period. Waves are sized for information rather than volume. The first wave’s job is to calibrate the stage conversion rates, so it deliberately mixes recipient types, channels, and work variants rather than concentrating on the likeliest targets. After a fixed window we read the funnel, find the one stage with a disproportionate drop, fix that stage, and send the next wave with the fix in place. We make one change per wave, because two simultaneous changes make the next reading unattributable. Across waves the targeting corrects itself as well. Stage rates broken down by recipient segment feed the next target list, so if one segment opens everything and never finishes, and another finishes everything it starts, the next list shifts toward the second segment. The system adjusts who it approaches as well as what it sends, both adjustments come out of the same event log, and we build the next wave’s work from what the silent majority observably did. **FIG. 3 · THE EVENT LOG** *Typed events arriving in the append-only log, each against one recipient. The funnel above is a query over these rows, and so is the next wave's target list.* ## Beyond outreach Nothing in this pattern is specific to outreach. Any agent-driven process that produces something and then waits for the world to respond has the same anatomy: a job posting waiting for applicants, an article waiting for readers, a support answer waiting for confirmation, a proposal waiting for a decision. In every case the outcome decomposes into stages, each stage can carry its own measurement, and the same two rules apply: every artifact gets its own token, so that events attach to one recipient rather than averaging over a campaign, and only one thing changes per wave, so that the next reading means something. A system without per-stage measurement can only vary its volume, and volume is the one adjustment that makes the channel worse for everyone who uses it. A system with per-stage measurement runs every wave as an experiment, and the silences, which are most of what any outreach ever hears, become the readings the next wave is built from. The world does not volunteer an error signal. A system that manufactures one out of its own silences can converge on outreach the same way it converges on code. --- # Encoding judgment as rules URL: https://idylliclabs.com/writing/encoding-judgment-as-rules Area: Governance · Published: 2026-07-17 Most of a person's judgment can be written down, and the written form is what governs. Corrections recorded as rules reach thousands of runs the person never sees, and every review makes the rules more complete. A person editing agent drafts rejects the same sentence shapes, the same layout mistakes, the same staging errors, week after week, and each rejection costs the same attention as the first one did, because nothing carries the correction forward to the next draft. AI agents produce drafts faster than any person can read them. That is the point of using them, and it turns the repeated correction into the system's bottleneck: the quality of anything crafted depends on judgment, the judgment lives in a person, and the person does not scale. Hiring more reviewers reproduces the problem the agents were supposed to remove, because now the reviewers' judgment has to be aligned too. What we learned building our systems is that the correction can be carried forward, if it is written down in a form every future run must read. ## From verdict to rule > The same discipline holds for documents aimed at humans: a rule stated abstractly is interpreted differently by every reader, and the worked example is what actually transmits. We run this as an explicit pipeline. A verdict is a binary judgment on one piece of work, with the reason stated: this draft fails, because the opening buries the strongest material. When the same verdict recurs, it is generalized into a rule, meaning a written instruction in a file that every future run must read before it starts. A rule carries annotated examples of passing and failing work, because agents follow examples more reliably than they follow abstract instructions; in practice the example is the specification, and a rule without one gets interpreted differently by every run that reads it. Where a rule is machine-checkable it additionally becomes a check, an automated audit that fails a draft on violation with no human present, so the rule is enforced at the moment of production instead of at review. The strongest form is a rule that becomes a structural fact: if a mistake was possible because a value was set by hand, the fix is to make the value derived, and then the mistake cannot be expressed at all, no matter how careless the run. The final property is inheritance. A run starting today reads every rule derived before it, and is checked by every audit written before it, so nothing has to be retaught. | FIG. 1 · FROM VERDICT TO RULE | | | --- | --- | | VERDICT | A binary judgment on one piece of work, with the reason stated. Recorded the session it is given. | | RULE | The verdict generalized into a written instruction that every future run reads before starting. | | EXAMPLE | Annotated passing and failing work attached to the rule. The example is the specification. | | CHECK | The machine-checkable form of the rule. Fails a draft automatically, with no human present. | | INHERITANCE | A run starting today reads every rule and is checked by every audit derived before it. Nothing is retaught. | *Read top to bottom: each row applies the judgment in the row above it more automatically. A verdict that stops at the first row has to be given again tomorrow.* ## A rule derived in three rounds The clearest worked example from our own operation is a writing register. One of our projects needed its public pages written in plain documentation prose, and the agents drafting them kept producing something else: dramatic fragments, metaphors, sentences arranged for rhythm. The rule took three rounds to derive. In the first round a person read the drafted copy and killed sentences one at a time, stating the reason each one died: this sentence narrates the page instead of describing the product, this one is a metaphor doing no work, this one exists for its sound. In the second round those verdicts were compiled into a rule file, built around a short corpus of prose in the target register for the runs to study, followed by a banned-pattern list with the verbatim verdicts attached to each entry. Drafts written under that rule were better and still failed in a subtler way, because the sentences were now compressed and manicured, performing plainness instead of being plain, and the verdict on that round became the rule's second revision. In the third round, drafts began passing on first review. The rule has cost nothing to apply since, and every page the project publishes inherits it. **FIG. 2 · WHAT THE ACCUMULATED RULES DO** *Twenty-six ways to draft the same work enter from the left. Each layer of recorded judgment ends some of them at its gate, and a path a gate has ended cannot be taken again.* Small judgments are recorded the same way as large ones. A reviewer once rejected a narration line because it counted the product's features, and the whole correction became one written sentence: never count features in a narration. The rule has held through every narration since, and nobody has had to give that verdict twice. Writing it down took under a minute, which is the general economics of the pipeline: a verdict is expensive because it needs a person's attention on real work, and everything downstream of it, the rule, the examples, the check, is clerical. The judgment is paid for once, and applying it afterward costs close to nothing. **FIG. 3 · RULES ACCRUING** *Each row cost one verdict on real work and nothing since. Every run that starts tomorrow reads the whole list.* ## Where the human ends up As rules accumulate, the person's role changes shape, and the change follows a fixed gradient. At the start the person approves each artifact before it ships, which is the correct posture for any new category of work, because the rules for it do not exist yet. When drafts start passing first review reliably, approval moves to batches: a whole wave reviewed at once. When batches pass, the person approves policies, meaning they approve the rules and the rules approve the work. The end state is review by exception, where only the low-confidence, unusual, or expensive cases reach a person at all. Each graduation is earned by the recorded evidence of the stage before it, and a category that starts failing again is demoted back down the same gradient. > Demotion matters as much as graduation. Without a way back down, a category would keep its autonomy after its rules stopped working. | FIG. 4 · THE APPROVAL GRADIENT | | | --- | --- | | APPROVE EACH | Every artifact is reviewed before it ships. The default for any new category of work. | | APPROVE BATCHES | Whole waves reviewed at once, after drafts pass first review reliably. | | APPROVE POLICIES | The person approves the rules, and the rules approve the work. | | REVIEW EXCEPTIONS | Only low-confidence, unusual, or expensive cases reach a person. | *Graduation between rows is earned by recorded evidence, and a category that starts failing is demoted the same way.* What remains for the human concentrates into two motions: approving, which is cheap, and correcting, which is the expensive judgment work. The discipline that makes the whole system compound is that every correction lands in a rule file in the same session it is given, because a correction that stays in the person's head applies to one draft, and a correction that lands in a rule applies to every draft after it. Judgment is usually described as the scarce input to AI systems, and usually as a property of gifted individuals. Our experience building these systems is that most of a person's judgment can be written down, and that the written form is what does the governing. The verdicts a person gives on real work, generalized into rules and enforced by checks, carry that judgment into thousands of runs the person never sees, and every review that produces a new correction makes the rules more complete. --- # Ideas with bodies URL: https://idylliclabs.com/writing/ideas-with-bodies Area: Persistent state · Published: 2026-07-17 An idea held in a head or a note cannot be found by a stranger, cannot take a payment, and cannot learn from silence. Agents have made it cheap to form the body an idea needs to meet its market, meaning an address, an inbox, an account, and records that persist. An idea is a pure informational object. It can be written down, explained to a friend, and believed with complete conviction, and none of that gives it a way to act. Held in a head or a note, an idea has no address, so a stranger with the exact problem it solves cannot find it. It has no channel, so it cannot ask anyone anything or answer anyone who asks. It has no account, so even a willing buyer has no way to pay it. Its only output is the feeling of having had it. The constraint was never the supply of ideas. Everyone with domain knowledge carries more ideas than they will ever test, because testing one has always meant months of a person’s attention: something has to be built, put in front of people, followed up on, and adjusted, and every one of those steps is labor. [The granular startup](https://idylliclabs.com/writing/the-granular-startup) describes the filters this price imposes on which markets ever get searched, and none of those filters selects for the quality of the idea. What is missing is not more ideas and not more willpower. It is structure that can carry an idea to its market without consuming its owner. Publishing the idea does not supply that structure. A post can be read, but it cannot follow up with a reader, quote a price, or notice that nobody came. Distribution, response, and adjustment were always the expensive part of the test, and writing the idea down moves it out of the head without providing any of them. ## Forming the body The capacities the idea is missing already have a description. [The agent body](https://idylliclabs.com/writing/the-agent-body) defines an AI agent as a mind plus a body, where the body is everything that persists between model calls and everything the agent acts through. An idea equipped with such a body has an address where strangers arrive, an inbox that receives messages whether or not anyone is attending, an account that can be paid, and records that keep what happened. The idea itself does not change. What changes is what can reach it. **FIG. 1 · THE SAME INBOUND FLOW, TWICE** *Strangers, replies, and orders rain on the same idea in both panels. As information it is a closed shell and everything deflects; as a body the boundary has gates, and the counter is what arrived through them.* Civilization already runs on a primitive for this. Incorporation means, literally, to form into a body: the law draws a circle around a set of accounts, assets, and obligations, and treats the circle as a person that can own, owe, and be dealt with. Forming such a body has been one of the expensive parts of testing an idea, because the entity, the bank account, the website, the email, and the books were each a professional service or a week of someone’s attention. These are exactly the parts agents can now stand up and operate, and Busibody, our infrastructure for businesses operated by AI agents, does this work: it handles incorporation, banking, payments, and email, and the owner approves anything consequential. With the setup and the operation handled by agents, the cost of giving an idea a body falls toward the cost of the compute. A generated landing page is not a body. AI can scaffold a site in an afternoon, but a site alone is still information, now hosted. The parts that make a body are the parts that persist and receive: - **The inbox** that holds a reply until someone reads it. - **The account** that can accept a payment. - **The record** of what was tried and what came back. The useful test is what remains reachable tomorrow: whether a reply sent tonight will be held, and whether an order placed next week can be taken. | FIG. 2 · THE SAME IDEA, TWICE | | | --- | --- | | AS INFORMATION | A one-line idea in a note. No address, no channel, no account. Nothing can reach it, and its only output is the feeling of having had it. | | AS A BODY | The same idea inside a boundary, with a site, an inbox, an account, records, and an approval gate attached. | | WHAT ARRIVES | Visitors, replies, orders, and silence. In the body, silence is information: a measured non-response to a real offer. | | WHAT EXITS | One channel, to the owner: anything consequential, meaning spending, contracts, and legal weight. | *The idea is unchanged. Everything that changed is what can now reach it, and what it can do about what reaches it.* ## The loop and the owner With the body formed, testing the idea becomes a loop: build, publish, reach out to the people the idea is meant for, read what comes back, adjust, and publish again. Agents run this loop for as long as the compute is paid for, and repetition is the part of a search that exhausts a person and costs a machine almost nothing. The person supplies the idea, the taste, and the judgment, and approves anything consequential, meaning spending, contracts, and anything with legal weight. The body messages its owner when a decision needs one. An idea a month can become a running experiment instead of a note, without the owner working more hours. A loop with no person inside it invites the objection that agents will flood every niche with generic work. The loop is not unattended, because the approval gate is a structural part of the body rather than a policy, and the question of whether the work converges toward quality or toward slop is answered in the granular startup’s terms: it depends on which verdicts from the market the loop is allowed to hear. We run businesses this way ourselves. One of them sells finished work: a customer places an order and receives a delivered piece of work rather than access to software. Agents produce the outreach and the work product, and the operating cadence is one question, asked of the business daily: is this business operational yet, and if not, what is needed. ## The verdicts that come back What the body sends back is different in kind from anything an untested idea produces. The owner of an inert idea learns nothing and risks nothing. Give the same idea a body and its owner starts receiving verdicts: an order, a reply, a refund request, or silence where a sale was expected. In the businesses we run, real outreach has gone out, and much of it has met silence or rejection. Those results are findings. Silence toward a real offer is information about a real market, and it is information the idea could not have produced from inside a note, where silence is indistinguishable from never having tried. A dead idea with a body has at least settled its question. **FIG. 3 · WHAT COMES BACK** *The verdicts a body collects, arriving as rows. Silence is one of them: a measured non-response to a real offer, which a note could never have produced.* The rejections carry a second value that is easy to miss. A rejection arrives addressed to the business, which has its own name and its own address, and not to the person whose idea it was. A test that carries its owner’s identity is expensive to fail in a way that has nothing to do with money, and that expense has kept many tests from running at all. Because the no lands on the business’s name, the owner reads it as data about an offer rather than a judgment of themselves, and in our experience more tests end up running. Delegating the outreach does not delegate the learning. Every verdict routes back to the person. What agents produce is the contact, meaning the outreach, the delivery, and the follow-up, and the reading of what comes back stays with the person. | FIG. 4 · THE ENGINE | | | --- | --- | | THE LOOP | Build, publish, reach out, read what comes back, adjust, publish again. Run by agents for as long as the compute is paid for. | | THE OWNER | Outside the loop, touching it at two points: a steering input into the idea, and an approval gate on spending, contracts, and legal weight. | | REPORTS BACK | Every verdict the loop collects routes to the owner. Delegating the contact does not delegate the learning. | | THE PORTFOLIO | The same loop, repeated. Several searches running at once, governed by the same judgment. | *The owner touches the loop at two points and stands inside none of it.* ## The portfolio If forming a body is cheap and the verdicts are cheap to collect, the natural unit of work changes. The same judgment can govern several of these searches at once, which the granular startup argues in its own terms. An idea then no longer needs to be the one bet its owner stakes years on. It can be one of several embodied ideas, each searching its own market, each cheap to wind down when the verdict is no. We run more than one such body ourselves, on infrastructure built so that each new body costs less to form than the one before it, and the direction is a portfolio: a person’s accumulated knowledge running as small businesses that refer work to each other. For each idea that gets a body, the owner ends up with a customer or an answer, and both are worth more than the feeling of having had the idea. Some of these businesses will settle their question and wind down. Some may come to pay their own way. Most, going by the verdicts so far, will hear more silence than orders, and the silences will be the first real facts most of these ideas have ever produced. --- # The granular startup URL: https://idylliclabs.com/writing/the-granular-startup Area: Economics · Published: 2026-07-17 Steve Blank called a startup a temporary organization searching for a repeatable business model, and nothing in the definition says how many people the search takes. When agents supply the labor the set of markets worth searching has to be recomputed, and at the limit the exploration of the markets startups ignore can itself be an automated process. Steve Blank defines a startup as a temporary organization built to search for a repeatable and scalable business model.[1] The definition is deliberately strange. A startup is not a small version of a company, because a company executes a model it already has, and a startup does not have one yet. A startup exists to run experiments until one of them finds product-market fit. Once one does, it stops being a startup. Paul Graham compares a startup to a mosquito: a bear can absorb a hit and a crab is armored against one, but a mosquito is built for exactly one thing, and everything that does not serve that one thing has been stripped away.[2] Both definitions describe a search process. Neither says anything about how many people it takes. The team, the office, and the payroll are not part of what a startup is. They are what the search happened to require, because until recently every experiment was made of human labor. Someone had to build the landing page, write the copy, run the outreach, answer the emails, read the replies, and adjust. The humans were the runtime the program ran on, and because the runtime never varied, everyone read the runtime as part of the definition. Taking the definitions at their word sets up an argument in three steps. A startup is a search program for product-market fit. The program is priced in its labor, and the labor has always been human time. When the labor becomes compute, the price of every experiment changes, and the set of markets worth searching has to be recomputed. ## The three filters Because the labor is human, the search is priced in human time, and human time is the most expensive input there is. Before anyone searches a market, three filters apply: - **The cost of finding out.** How many months of a person’s attention it takes to learn whether demand exists at all. - **Margins.** Whether the niche, if it works, pays enough to keep repaying that attention. - **Longevity.** How long the niche survives before it is competed away or made obsolete. A market has to clear all three before a person will commit a year of their life to it. **FIG. 1 · THE THREE FILTERS, APPLIED** *Four hundred niches with real demand enter at the top. The dots falling between rows are the markets each filter removes, and the residue at the bottom is where startups have always searched.* Most niches fail at least one filter. Consider a product that a few hundred people would pay for, worth perhaps two thousand dollars a month at its peak, in a niche that a platform change will erase within a quarter. The demand is real and the money is real, but the same year of a person’s attention could be spent on a market a hundred times larger, so nobody searches this one. Niches like this are cracks in the market, and value seeps through them continuously. Every niche too small to repay a person, too thin in its margins, or too short-lived is left on the table, and there are far more markets below the threshold than above it. None of this is a fact about the markets. It is a fact about the price of the search. ## The constant changed Agents changed the price of the experiment. An agent can build the page, write the copy, run the outreach, read the replies, and adjust, and it can repeat that loop for as long as the compute is paid for. Trying again is exhausting for a person and close to free for a machine. So if a startup is a temporary structure built to search, what is the smallest structure that can still run the search when agents supply the labor? Our answer is the granular startup: a business that is 80 to 90 percent agent-directed, with a human governor. Agents run the experiments end to end, meaning they build, publish, sell, answer, and adjust. The governor sets the direction, supplies the taste, and approves anything consequential, such as spending, contracts, and anything with legal weight. The structure abstracts away the effort and the mental toil, and the judgment stays with the person. At the new price, the three filters read differently: - **The cost of finding out collapses.** The experiments are made of compute, and a granular startup can probe a market for less than it used to cost to think seriously about probing it. - **Thin margins clear.** The structure’s operating cost is a fraction of a salary. - **Longevity stops being a filter at all.** A niche that will be driven out of existence in three weeks is still worth entering, because standing the business up and winding it down both cost almost nothing. The band of markets that sit between what repays an agent’s time and what repays a person’s time is exactly the value that has been seeping through the cracks, and the granular startup is the structure that collects it. | FIG. 2 · THE TWO THRESHOLDS | | | --- | --- | | REPAYS A PERSON | Markets large, durable, and rich enough to return a year of a person’s attention. The few at the head of the tail, where startups have always searched. | | BETWEEN THE LINES | Markets that clear an agent-directed search but not a person’s: real demand, thin margins, short lives. The value seeping through the cracks. | | BELOW BOTH | Markets too small to repay even compute. Left alone by both kinds of search. | *The middle band is wider than the top one, because there are far more markets below a person's threshold than above it.* **FIG. 3 · A FIELD OF MARKETS, TWO SIZES OF SEARCHER** *The dashed circle is a person's search, larger than every small cell, so it can only land on the few large markets. The moving dot is an agent-directed search entering the small ones, one cheap experiment at a time.* ## The slop objection The obvious objection is slop. Agent businesses can produce slop today, and an agent left alone with its own output will confidently produce generic copy, plausible-sounding answers, and a mediocre product. The mechanism is that a model iterating on its own output updates only on internal consistency. Each pass gets more elaborate and no more accurate, because nothing outside the loop pushes back. The correction is reality contact. When the loop includes the world, meaning a buyer who pays or does not, a refund request, a reply, a complaint, or silence where a sale was expected, each iteration updates the business toward what the market actually wants. This is what the agent loop is good at: it is a search algorithm with a convergent step, and it will run the step as many times as it takes. An agent business with enough reality contact and intelligent corrective loops converges away from slop the way any feedback system converges. We do not claim this is solved. We claim it is the design problem: deciding which verdicts from the market reach the agents, how quickly they arrive, and what the agents are permitted to change in response. Busibody, our infrastructure for these businesses, is largely an attempt to engineer that convergence: incorporation, banking, payments, and email exist in it so that real verdicts can flow in and consequential actions can flow out under an owner’s approval. | FIG. 4 · CONVERGENCE WITH AND WITHOUT REALITY CONTACT | | | --- | --- | | LOOP CLOSED ON ITSELF | Each pass updates on internal consistency only. The work gets more elaborate and no more accurate, however long it runs. | | LOOP INCLUDING THE MARKET | A sale, a refund, a complaint, silence where a sale was expected. Each verdict moves the business toward what the market wants. | *Both loops run on the same model and the same compute. What differs is what is allowed to push back.* ## The limit None of this competes with human entrepreneurship. A granular startup keeps a human at the top for the same reason it exists at all: the scarce ingredient was never judgment. There are far more people with sound judgment about some corner of the world, a trade they know or a community they belong to or a problem they have watched go unsolved for years, than there are people who can afford to spend a year of labor testing what they know. > The three filters never selected for the best judgment. They selected for whoever could pay the search cost. When machines supply the labor, the judgment that was always there finally gets to run its experiments. It does not have to run them one at a time. The structure is cheap enough that a governor can direct many granular startups at once, so the same judgment can be searching dozens of markets in parallel, each search collecting its own verdicts. The structure also removes a familiar inflexibility. A traditional startup that concludes its first product is not working pivots, and a pivot replaces the whole identity. When Wordware, a workflow-builder company, moved to the AI companion Sauna in 2025, the change meant a complete architectural rebuild and a team realignment, a stretch the company itself called the hardest period in its life.[3] The pivot costs that much because the company is one body with one identity, so trying a different business means becoming a different company. A granular startup forks instead. The same governor runs several versions of the business as simultaneous experiments, pooling the same resources and the same users, and keeps whichever version the verdicts favor. No version has to die for another one to be tried. The price of that search is not done falling, because every part of a granular startup that is made of compute gets cheaper on the model vendors’ schedule, not ours. At the limit, the last human bottleneck comes into view: choosing which market to probe next. Choosing is itself a loop. It reads the field, proposes a candidate niche, stands up the smallest structure that can test it, reads the verdicts, and keeps the business or winds it down. That is the same search one level up, run across markets instead of within one, and nothing in it requires the person to be more than the governor they already are. That is why the granular startup is an existence proof rather than a product category. One of them proves the unit: a structure this small can find a market, serve it, and collect value that was seeping through the cracks. A process that spawns them explores the space. There are vastly more niches with real demand than there are people positioned to search them, and the filters that kept it that way were facts about the price of labor, not facts about the markets. The price changed. There should be much more of this. REFERENCES - [1] Steve Blank, [“What’s A Startup? First Principles.”](https://steveblank.com/2010/01/25/whats-a-startup-first-principles/) steveblank.com, 2010. - [2] Paul Graham, [“How to Make Wealth.”](https://paulgraham.com/wealth.html) paulgraham.com, 2004. - [3] Filip Kozera, [“Filip Kozera, CEO of Wordware, on the rise of vibe doing.”](https://sacra.com/research/filip-kozera-wordware-rise-of-vibe-doing/) Sacra, 2026; and the company’s own account at wordware.ai/story.