Report · Wednesday 2 September 2026
You approved a five-dollar test to find out what one of these agents really costs. It is built, it has a hard cap on it, and it is one command from running. While building it I measured the thing my earlier estimate was guessing at — and the honest answer is that I was probably out by a factor of ten, in your favour. This report says what I measured, what I am still guessing, what the four agents would be, and the one I would not build.
Read this first Still nothing has been spent. No agent has been created, no session started, no money has left the account. The test is built and waiting on a terminal that holds the account key — this one does not have one.
Every cost figure in this report is modelled except one, and the one that is measured is clearly marked. I would rather hand you a number with its workings attached than a confident wrong one.
The checker is built and read-only. The one you asked for yourself — "some kind of AI quality control that says it did it." It cannot send, post or spend, and that is not a promise this time. It is thirteen checks in the test suite that fail if anyone changes it.
My five-dollar estimate was probably ten times too high. I guessed it from published rates and a guessed workload. I have since counted the actual workload. Same rates, real input, and the checker comes to about thirty cents per client per month — roughly fifteen dollars at fifty clients, not two hundred and fifty.
It is still an estimate. The input is measured. How many turns the agent takes and how much it writes are still modelled. The run settles it, and that is exactly what the five dollars is for.
All four agents together model at about $1.28 per client per month. That is about $64 at fifty clients and $320 at 250 — the shape you are aiming at. Less than one of the two-hundred-dollar accounts.
The tier you asked for got cheaper while we were not looking. You said "probably Sonnet 4.6 at least." Sonnet 5 is the newer model and the cheaper one. Asking for 4.6 today would cost more and get less.
Chad is a real client with real callers. This was the first thing I checked, not the last.
Yes — and not because I am being careful. Because there is nothing there to be careless with.
| What it cannot do | Why not |
|---|---|
| Reach the internet | Its work space is built with an empty list of allowed sites. Not "we told it not to" — there is no way out |
| Touch the CRM, email or phones | Nothing is connected to it. There is no wire to GoHighLevel |
| Search the web | Switched off on the agent itself. This one matters: web search runs on Anthropic's computers, not in our work space, so blocking the work space would not have stopped it |
| Run commands, or edit anything | Both switched off |
| Keep anything | Its whole world is three read-only files and a scratch disk that is destroyed when it finishes |
The only thing it can write is its own report.
And those five lines are held by a test, not by this report If someone turns the shell back on, or opens the network, or raises the five-dollar cap, the test suite fails and names which promise broke. I checked it actually works by deliberately breaking two of them in a copy.
We have been here before. A switch that would have texted a homeowner sat switched off for weeks under a comment saying nothing there texts the customer. A promise with nothing holding it is the same shape of problem.
Short answer: the Checker touches neither, and three of the four never would. Here is how the data actually gets to it.
First, so nobody goes looking for something that is not there Your Anthropic console is blank, and that is correct. Nothing has been created in the account. The Checker is written, not created — it comes into existence the first time we run the test, and not before.
The instinct is that an agent must reach out and fetch what it needs, and that we would therefore have to open a door in our server or hand over a GitHub key. The Checker works the other way round. We push the data to it.
| Step | Where it happens |
|---|---|
| 1. Pull the client's month — calls, reader findings, their brain | Our machine. Straight out of our own data |
| 2. Write it to three small files | Our machine |
| 3. Hand those three files to Anthropic | We upload them. The agent does not fetch anything |
| 4. The agent reads them and writes its report | Anthropic's work space. Files mounted read-only |
| 5. We collect the report; the work space is destroyed | Nothing is kept |
So the door is never opened. The agent has no way to reach out — its list of allowed sites is empty. It cannot pull from us because it cannot reach us. It only ever sees the three files we chose to hand it.
| Agent | Rose Hosting | GitHub | What it does need |
|---|---|---|---|
| The Checker | Never | Never | Nothing. Three files we push in |
| The Brain Keeper | Never | Never | Same — we fetch the live setting and hand it over as a file |
| The Newsletter Writer | Never | Never | Web search, which runs on Anthropic's computers, not on ours |
| The Microsite Keeper | Never | Yes — to propose a page change | A key limited to one repository |
Not one of the four ever touches Rose Hosting. That box runs Josh, the phone line, the booking system and the dashboards, and nothing here goes near it. It also does not go away — this is extra, and it replaces nothing.
The one that does need GitHub, and why it is still safe Only the Microsite Keeper, and only because proposing a page change means opening a pull request. Two things about that key:
It is limited to one repository, not the account. And it is never put inside the work space. Anthropic holds it outside and adds it to the request on the way out, after it has left. Nothing running in that work space — including anything the agent itself writes — can read it or take a copy.
And it still cannot publish. A pull request is a proposal. A person merges it.
I pointed the measuring tool at our own data. Chad's last thirty days:
| What | |
|---|---|
| Calls on his line | 434, of which 233 booked |
| Calls the transcript reader has been through | 108 |
| His onboarding brain | The draft record — which is what a caller actually hears |
| All of it, as text | 171,102 characters |
Chad is our busiest line by a distance — 434 calls against Quality's 103. So this is a ceiling for one client's month, not an average.
| My 1 Sept report | This report | What changed | |
|---|---|---|---|
| One client, one month | ~$5.00 | ~$0.30 | The old number multiplied real rates by a guessed workload. This one multiplies them by a counted workload |
| Fifty clients | ~$250 | ~$15 | Same |
Do not take thirty cents as fact either Here is exactly what is solid and what is not:
Measured: the 171,102 characters, counted from our files today.
Published: the rates — $2 and $10 per million words-ish, $10 per thousand web
searches, 8¢ an hour of work space. Checked against Anthropic's own page today.
Still guessed: how many times the agent goes round the loop, how much it writes,
and how much of its reading is served cheaply from memory.
The run replaces those last three. That is the entire point of the five dollars.
Where the money goes on one run, so you can see there is no hidden line:
| Part of the job | Cost |
|---|---|
| Reading the evidence, first time | $0.0960 |
| Putting it in memory | $0.1125 |
| Re-reading it from memory | $0.0300 |
| Writing the report | $0.0600 |
| Web searches | $0.0000 |
| Work space time — four minutes at 8¢ an hour | $0.0053 |
| Total | $0.3038 |
Look at the last line before the total. Four minutes of a computer costs half a penny. That is the "we need no new server" claim from last week, written as money instead of as an assurance. And to say it again plainly: this replaces nothing we run today. Josh, the phone line, the booking system and both dashboards still need Rose Hosting.
Dil asked this, and it is the right question. The honest answer is that it does not fix Josh — it finds what nobody is reading.
Let me be straight about this rather than talk around it. Fixing a Josh fault takes five steps, and an agent with no reach into our server can do one of them:
| Step | Who does it |
|---|---|
| 1. Notice the fault | An agent can do this |
| 2. Work out why | Claude Code and a person |
| 3. Change the code or the client's record | Claude Code and a person |
| 4. Ship it — merge, then the deploy button | A person answers yes. The restart drops whatever is live on Chad's and Ted's lines |
| 5. Prove it on a real call | A person rings it. An agent can read the result |
So it is a smoke detector, not a fire engine. Nothing here changes how a fix gets made or shipped.
But look at where Josh's faults actually survive They do not survive because we could not fix them. They survive because nobody read the thing that already contained them. Three counts, all from our own files:
Nine of the eighteen faults from the 28 August calls were already in Hamming's own transcripts — it produced the condition, recorded the fault, and graded the call a pass, because no assertion looks at that. The evidence was sitting there.
Forty of sixty-eight test calls have the caller asking Josh if he is still there. Six calls in ten. None of it graded.
Of 825 real calls in the log, 147 have ever been read — 18%. Of those 147, 31 turned up something. If that rate holds across the 678 nobody has read, there are roughly 140 findings sitting unexamined. That last figure is an extrapolation, not a count.
And the fault that cost a real job — Josh refusing a pre-drywall quote Chad sells at $425 — has this written next to it in our own notes: "It cost a job on a real call and nothing flagged it."
That is the gap. Not fixing. Reading. And reading a large pile cheaply and every single month is the one thing these agents are unambiguously good at.
One correction before anyone buys an agent to solve this For the nine faults Hamming already recorded, the cheapest fix is not an agent. It is the missing assertion. Our own analysis prices one of them at "free — one assertion". Writing the assertion catches that fault every time, forever, at no running cost.
The agent is the backstop for everything no assertion was ever written for — which is most of the real calls, because real callers do things nobody thought to test. Do both, in that order. Assertions first, because they are free.
Josh's brain is not only code. A large group of his faults are the client's own record being wrong — a price missing, a service with no figure, hours that do not match the days he books. Four of the eighteen August faults were exactly that: the client's card, not Josh.
Those do not need a code change or a deploy. They need one person to type one value into one record. The pre-drywall fix was a number reaching Chad's onboarding record — the earlier attempt went into the code and callers never saw it.
That is what the Brain Keeper is for, and on that class of fault the distance from "found" to "fixed" is a minute, not a release.
You asked what they would be, what they would do, and how we would train them. Here they are, in the order I would build them.
The bar for each is the one we set when we archived sixteen specialists in August: a real job is waiting, and the agent gets the tools that job actually needs. Those sixteen failed the second half — every one had the same toolset and none of them could reach the thing they were named after.
What it does. Once a month, per client: reads their calls, the transcript reader's findings, and their onboarding brain. Answers one question — did the service we sell actually happen — and names every place it did not. Then checks the brain itself for anything that would make Josh answer wrongly: a service with no price, two prices that contradict, hours that do not match the days he books.
How we train it. We do not. We write it — a page of instructions naming the five sections of its report and the rules it works under. Every claim must cite the call it came from. Never quote a Booked Solid price. Say "I cannot tell from this" rather than fill a gap. Changing it means editing that page.
Why it is not busywork. Three of the worst faults we have had were exactly what this looks for, and all three sat unnoticed: Ted's Delaware line answering with Chad's Texas prices, the pre-drywall quote Josh refused on a real call, and a number that looked perfectly mapped and played silence. Nothing flagged any of them. At fifty clients nobody is checking by hand.
What it does. Reads what a client's live setup actually says — not the version in the code, the live one — beside their onboarding answers and a sample of what Josh really said on calls. Flags where the three disagree. Monthly, and again whenever an onboarding form changes.
How we train it. One rule, which is most of the job: an unfinished draft form beats a submitted one, which beats the built-in default. A correction typed into the code is invisible on the phone until it reaches the record. That single sentence is why the pre-drywall fix took a week to actually land.
Why I am proposing it. It is the cheapest of the four and it guards the fault that has cost us most. Our own note says the mechanism behind that fault "will make the next one possible too." This is the thing that reads the live brain so a person does not have to remember to.
What it does. Once a month, per client: finds the content, writes the homeowner letter and the agent letter — two different readers, two different sources — and a short note saying what it used. It drafts. It does not send.
How we train it. Three things, and the first is yours. The twelve client emails and twelve agent emails in your voice. Then the client's own brain, so a Nashville inspector's agent letter can be about the Nashville market — which you asked about and nobody has decided. Then a shared notebook of what last month said, so December does not repeat November.
The honest bit. The monthly timer is small work — the scheduler already runs timed jobs and has simply never been asked to repeat one. It is blocked on the words, not on us. An agent given no voice will invent one, and it will be nobody's.
What it does. Keeps the town pages fresh every month — the local detail that makes a microsite worth having. It proposes the change. It does not publish.
How we train it. A page template and a rule about what counts as local. It writes into the shape; it does not invent the shape.
And the important half. Build the microsites themselves in Claude Code, free. Last week's report measured this: our own 31-page site is about 1.27 million characters, nothing extra our usual way, forty to ninety dollars an agent's way. The agent is for the upkeep, which comes back every month. The build comes once.
Fifty by year end and 250 in year one — costed at the shape you are actually aiming at, not at today's handful.
| Agent | Per client | 2 clients | 50 clients | 250 clients |
|---|---|---|---|---|
| The Checker | $0.30 | $0.61 | $15.19 | $75.96 |
| The Brain Keeper | $0.18 | $0.36 | $9.10 | $45.50 |
| The Newsletter Writer | $0.44 | $0.89 | $22.15 | $110.75 |
| The Microsite Keeper | $0.35 | $0.70 | $17.48 | $87.38 |
| All four | $1.28 | $2.56 | $63.92 | $319.58 |
And the model choice, priced rather than argued:
| Model | All four, one client | At fifty | |
|---|---|---|---|
| Sonnet 4.6 | $1.82 | $91.24 | the tier you named in the meeting |
| Sonnet 5 | $1.28 | $63.92 | newer and cheaper than the line above |
| Opus 5 | $2.92 | $145.89 | the expensive one; not needed here |
Sonnet 5 for all four. These are checking and writing jobs where the hard part is being thorough, not being clever.
Your words: "I just don't know how it did it." So here is how, rather than that it does.
You give it a time — every Friday at 8pm, say. At that moment the system creates a fresh work space, the agent does the job, and the work space is destroyed with it.
Nothing sits running in between, because there is nothing to sit running. And this is not tidiness, it is how the bill works: work space time is charged at eight cents an hour, and only while a job is actually running.
Four minutes of the Checker costs half a penny of work space. A month of a server doing nothing costs a server. That is the whole difference — and the reason this needs no new box while changing nothing about the boxes we already have.
The test prints the seconds it was actually running and then deletes itself, so you can watch this happen rather than take my word for it.
It is more literal than "shared notebook" sounds. It is a folder.
The folder lives outside any one job. When an agent starts work, the folder appears inside its work space, and it reads and writes it with ordinary file commands. Four things matter to you:
It survives. The job is temporary; the folder is not. Next month's newsletter
agent opens the same folder this month's wrote to.
It is shared, not copied. The Checker and the newsletter writer open the
same folder. Nobody carries a message between them — which is your "the
playbook agents aren't talking to the outcrop agents" complaint, answered by there
being one folder instead of two.
Every change is kept. You can see what changed, when, and undo it. That is the
real answer to "they go stale" — a stale note is visible and removable, not a
personality that drifted.
Reading and writing are separate permissions. The Checker gets read-only. It can
learn from what others wrote; it cannot rewrite the record.
Last week's report was willing to say two of seven accelerants were a bad fit. Same discipline here.
| Not building | Why not |
|---|---|
| Paid Ads | The technology fits — it could read the numbers, write the report, draft the changes and stop before the money. But there is no ad account spending anything for it to read. An ads agent with no ad spend is a charter, not a worker, which is exactly what got sixteen specialists archived. It passes the day there is spend. Not before. |
| Authority Builder | A one-time job, not something that comes back every month. Still unruled by you, too |
| The Converting Website | Build it in Claude Code, free. We measured that last week |
| Any kind of orchestrator | Your words: "the orchestrators got in the way… we had people sitting that we weren't using… and they got rogue, and they got stale." These four report to nothing. Each wakes on its own, does one job, writes to the folder, stops. If one is not earning its keep you delete it and nothing else notices |
The free Google Business Profile offer is still live in five places Including the public website and Parker's script. You closed that offer. Until it is changed, Parker will promise it to the next inspector who asks — a promise made in our name that we do not intend to keep.
Last week's report asked for your word on this and has not had it. It reaches customers, so no agent touches it without you. One sentence from you unblocks it.
In this order
What only you can do