Daily report · Saturday 5 September 2026
Two teams worked on Josh today. One built a machine that tests him every two hours, day and night, and writes the fixes while everyone is asleep. The other worked straight down Chad's list of complaints. Six of his nine are fixed and on the phone now. One is not fixed, and we say so. Two need a decision from Ken. Nothing was sent to a client. Part one is the machine. Part two is Chad's list. You can read either one on its own. What still needs you is at the end.
Read this first Nothing reached a client today. Everything we changed is code and tests in the repo, plus settings in Hamming and GitHub. The phone fixes went on the live phone when Dil pressed the deploy button at the end of the day. It worked.
The biggest thing we found was not new. Our own call checker had already flagged 21 real calls in two days, and 13 of them were serious. Nobody was reading them. The same four mistakes kept showing up. That is the "sometimes right, sometimes wrong" you have been hearing, with Josh's own words to prove it.
Why this report is long Two work sessions ran today, and both were busy. They are kept apart on purpose. Part one is the testing machine Dil asked for this morning. Part two is Chad's list of things Josh gets wrong on the phone. If you only read one, read part two. It is the list you gave us, and six of the nine are done.
A report writes itself every two hours. It reads every real call and every note the checker wrote about it. It looks at every Hamming test for silence and stalls. It replays Josh's safety rules over 101 recorded calls. And it runs all 219 tests. One word at the top tells you what to do: QUIET, ATTENTION, or REGRESSION. A REGRESSION opens a GitHub issue, so it reaches a phone.
Hamming calls Josh on a timer. Two "heartbeats" are set up, one on Chad's line and one on Ted's. Nine test runs a day, never both lines at once. Dil used to upload a CSV and forward the results by hand. That step is gone. The results go straight into the report.
A Claude session writes the fixes at night. It reads the newest report, writes the fixes with tests, and leaves Dil a patch file and a short note. Dil's morning job is: read the note, apply, merge, press deploy. About two minutes.
Three Josh mistakes are fixed in code, not in the prompt. Josh saying "your inspection is Friday" right after saying the booking failed. Josh offering "this afternoon" when the diary had nothing until Thursday. And Josh's own instructions telling him to say "email/text" — the one word your no-texting rule bans.
Six things on Chad's list are fixed. Josh no longer asks if the power is on when someone lives in the house. He no longer asks you to agree to the 5-Star after you already said yes. He says "or for someone else" instead of "or for something else". He now understands an email spelled out in code words. Tomball and The Woodlands are covered. And he leads with the 5-Star instead of offering a choice.
Every fix was counted first, not guessed. We have 110 recorded calls. Before writing a fix we count how often the problem shows up in them. Chad's utilities complaint was real 20 times. The 5-Star one was real 8 times. That way we know a fix is worth making, and we can prove it worked.
You decided on the contract. We have four early adopters, and clients do not know what to do in ISN. So every agreement, email and thank-you letter comes from Booked Solid as a template, and Brain 2 decides what is different for each client. A standard contract is drafted from Chad's and Ted's. It needs a lawyer to read it and four new boxes on the sign-up form. Then the e-sign can go on.
The list of tasks, in the order we did them.
| Task | Result |
|---|---|
| Read the whole Josh setup and the last three days of call logs. Found the 21 flagged calls nobody had read. | We know why Josh is "sometimes wrong", with quotes |
| Built the every-two-hours report and its tests | PR #630, merged |
| Ran the first report on GitHub. Found and fixed two bugs in the report itself. | PR #632, merged |
| Wrote three safety rules for Josh and checked them against 6,917 recorded sentences | In PR #632, merged |
| Set up the night-time Claude fixer | Runs every night and every morning |
| Set up Hamming heartbeats on both client lines. First GC run passed 6 of 6. Quality failed one case. | Live, nine runs a day |
| Dug out why the e-sign was switched off. Tested the contract sender against both clients' records. | See item 7 |
| Meeting with Ken and Beth: four early adopters, and the template decision | Decided |
| Drafted the standard contract. Wrote today's decisions into the repo so the next Claude session knows them. Wrote this report. | PR #637, merged |
| Second session worked through Chad's nine complaints | PRs #634, #635, #636, #638, #640, merged |
| Added the two towns Chad asked for to his record | Live now, no deploy needed |
| Put all the phone fixes on the live phone | Dil pressed deploy. It worked |
Six items. Each one says what it does and why.
Dil's words this morning: "When I'm asleep no one is continuing to test and fix."
A job on GitHub's own computers now runs every two hours, day and night. It needs no server, no keys, and no person. It does four things. It reads every real call Josh took and every note the call checker wrote. It scores every Hamming test for two things the log said to measure but nobody did: silent calls, and how often Josh says "one moment, I'm still here." It replays Josh's safety rules over 101 recorded calls and compares with last time, so if a fix makes Josh drop a sentence he used to say, we find out. And it runs all 219 tests, sorted into "was already failing" and "just started failing".
One word at the top. QUIET means nothing new. ATTENTION means some real calls went off track and it is worth a read. REGRESSION means something is worse than last time, and it opens a GitHub issue so it reaches a phone.
It cannot change the live code and cannot touch the server. A test makes sure of that. Changing the live code restarts the app that answers Chad's and Ted's phones, and this job runs while those phones are live.
The call checker you asked for on 24 August has been doing its job. Nobody was reading it. Twenty-one real calls in two days, thirteen of them serious. The same mistakes keep coming back:
Josh confirms a booking right after saying it failed. On a Staffordshire call: "I can't finish the booking on my end just now" and then, right away, "Your inspection is Friday, September 4th at nine in the morning with Glen Williams." Nothing was saved. A Quality caller heard the same thing: "we'll see you this afternoon."
Josh offers a time the diary never had. Quality's diary had nothing until Thursday. Josh: "We have an opening this afternoon at three o'clock with Ted Hinderer." Staffordshire needs two days' notice. Josh: "We have an opening today at two PM."
Josh guesses a price and books without checking. Twice, on two clients. He picked the price by guessing instead of using the price tool.
Josh tells callers about an e-sign agreement and a payment link. Four callers, two clients. No client asked for this, and the e-sign is switched off. Why it was off, and what you decided, is item 7.
Your rule from yesterday: "a prompt line is only a strong suggestion." All three fixes follow it. Each one was checked against every sentence Josh has ever said on a recorded call, 6,917 of them, both ways: does it catch the mistake, and does it leave the good sentences alone.
Saying a day counts as saying "you're booked". "Your inspection is Friday" and "we'll see you this afternoon" are now blocked when nothing has been saved. Josh says the honest line instead. Out of all 6,917 sentences, this touches exactly one, and that one came after a real booking, where the rule is off anyway.
A time the diary never offered is blocked. If Josh offers a day the diary did not give him and the caller did not ask for, the sentence is swapped for "let me check exactly what we have open." Reading back, the recap, "is Friday good for you?" and the confirmation after a save are left alone. We proved this inside the real call flow, not just on its own, because four rules before this passed their own tests and then never fired on a real call.
The closing line no longer says "email/text". Both of Josh's instruction sets were handing him the word "text" in the very sentence your no-texting rule was built to catch. Now it says "email, never a text", in both files.
Dil's words: "Usually it was me putting a CSV in the Hamming, then once there's the JSON result we forward it to Claude Code."
Hamming has a feature called a heartbeat: pick the agent, pick the test cases, set a clock. Two are now live. Josh GC calls Chad's line six times a day. Josh Quality calls Ted's line three times a day. Six test cases each, never both at once. Nine calls at once is what crashed the phone on Thursday, and six never has.
The results come to us on their own. Dil added the Hamming key to the repo this morning. Every finished test run is pulled into the two-hour report and scored. No download, no forwarding.
First results. GC's first run passed six of six. Quality's first run failed one case. That failure is already in the next report, and it is the first thing the night fixer will read.
Every evening and every early morning, a Claude session wakes up on its own. It reads the newest report and the checker's notes, picks the mistakes that keep repeating and that code is allowed to fix, and writes the fix with tests, following the same rules as item 3. It leaves Dil a patch file and a short note in plain English. If there was nothing worth fixing, it says so in two sentences and stops.
Dil's morning job takes about two minutes. Read the note. Drop the patch into a Claude Code session. Say "apply, push, open a pull request." Merge. Press the deploy button.
Nobody fixes the live phone at night, on purpose. The fix is written and tested while everyone sleeps. It goes live when a person presses the button, because that restart hangs up on anyone who is mid-call with Chad or Ted. That is your rule, and it stays.
The night session never changes the live code, never deploys, never starts a Hamming run, never spends money, and skips anything marked as your decision.
The first report saved the whole repo and no report. The publishing script named a folder that did not exist yet, and a backup step saved 1,858 files instead. Fixed. The test now runs the real script three times on a throwaway copy and counts what lands. The bloated branch cleans itself up on the next run.
The first run cried wolf. One test passes on a laptop but fails on GitHub's machines, because it uses a lot of memory on purpose. The report called it a regression and opened an issue. It is now on the "known to fail" list with the reason. That issue can be closed.
The weekday schedule was set for American nights. That is daytime in the Philippines, the one time Dil did not need it. Changed to every two hours, around the clock. That patch is in Dil's hands.
Too much clicking. The session Dil worked in could read the repo but not write to it, so every change had to travel as a file through a second session that could. That is a settings difference between two kinds of Claude session, not a limit on what Claude can do. Next time the work starts from Claude Code with the repo attached, and the file-passing goes away.
Item 7. The e-sign, why it was off, and the standard contract.
Your words in the meeting with Beth: clients do not know what to do in ISN, so every agreement, contract, email sequence and thank-you letter must come from Booked Solid as a template, and Brain 2 decides how different they are.
The emails were already done. You made them one standard set on 3 September. The contract was the last thing still made per client, and the code still followed the old rule from 9 August that Chad's and Ted's must stay different.
Why the e-sign was off. Nobody decided to keep it off. It was built and hooked up on 11 August. A check-up on 14 August saw the switch was off and wrote "Ken's or Dil's call". Nobody asked. So for three weeks Josh promised an e-sign that never got sent. Today we also tested the contract sender against both clients' records: Chad's bookings would send. Greg's would be refused, because he has no licence number on file. Every Quality booking would be refused, because Ted's record has no time zone and no licence numbers for Ted or Joshua. Those have been missing since 9 August.
A standard contract is drafted. Built from Chad's and Ted's word for word. No legal sentence was made up. Everything the two disagree on became a blank that Brain 2 fills in: the state, the licence, which rules of practice, the liability cap (Chad's is the fee, Ted's is twice the fee), who settles disputes, the return-visit fee, and how long to make a claim. Texas and Delaware each get a short block of their own clauses. A state with no block is refused, so no one is ever handed another state's law.
What this needs from you before the switch goes on A lawyer reads the standard contract once for each state we offer it in. Texas and Delaware today. This is the paper that gets your clients paid, so it is worth one read each. Chad's and Ted's own contracts stay in place until that is done.
Four boxes on the sign-up form: state, each inspector's licence number, the liability cap, and the return-inspection fee. Right now we are missing Greg's, Ted's and Joshua's licence numbers and Quality's time zone.
The $1 test payment link comes off once Chad's Stripe is connected in his own account. Otherwise a real contract goes out next to a $1 bill, which would confuse a customer.
His nine complaints, plus two things worth knowing that we learned while working on them. Each one says what he said, what we found, and what we did. His full list is in yesterday's report.
How each of these was checked We have 110 recorded calls. We count the fault in them before we fix it. That does three things. It tells us the complaint is real. It tells us how big it is. And after the fix we can count again and see it gone. Nothing here is a guess.
He is right, and it happened a lot. In our 110 recorded calls Josh asked it 20 times, on 13 different calls. Each time the caller had just told him someone lives in the house.
Fixed. If someone lives there, Josh does not ask. The record still says the power is on, and the log says we worked that out rather than asked. So the inspector still sees the field filled in.
The easy version of this fix would have been dangerous. Josh has to write down "someone lives here" or "the house is empty". We tested eleven real ways people say it. Five of them were being written down wrong — "it's vacant", "the house is empty", "nobody lives there", "it's unoccupied", and just the word "Vacant". All five were recorded as someone lives here.
So the quick fix would have stopped Josh asking about the power on the empty house — the one house where the answer decides whether the job can happen. We only found that because we counted first. It is fixed too.
We found the call he means. Here it is, word for word:
Josh: Want me to set you up with the five star?
Caller: Yes, that sounds good.
Josh: The five star package is $572. Does that work for you?
Caller: Yes, that works for me.
Josh: The five star package is $572. Does that work for you?
Caller: Yes, that works for me.
Three yeses to one question. Every extra ask tells a customer you did not believe the first one.
Fixed. Josh can still say the price. He can still say what is in the package. He can still read it back at the end. What he cannot do is ask you to decide it again.
The tricky part was one sentence meaning two things. "Does that work for you?" after a price is a question the caller answered already. The same words after a time are a real question about the day. Josh now knows which one he is asking. We checked: without that, the fix would have eaten three real questions about the appointment day.
Ken's answer: "we should say someone else." A person is not a something.
The right words were already written down. The question we wrote for Josh reads: "is this inspection for you as the buyer, or are you booking for someone else?" Nothing ever handed it to him. What he actually got each turn was a list of three words: buyer / homeowner / buyer's agent. A list of three words is not a question, so he made one up. The one he made up is the one Chad heard.
Fixed for every question, not just this one. Josh is now handed the exact words for each thing he has to ask. That is a code change, not a note in his instructions, so it happens every time.
A caller spelled an email out in code words — "D for Delta, O for Oscar". Josh read the code words back as if they were the address.
Same story as the last one. The code that turns "D for Delta" into the letter D has been in the system since 3 September. It was only plugged into a watcher that noticed the problem. Nothing gave the letters to Josh. Now it does: he is handed spelled so far: DORTIZ and reads that back.
Chad heard Josh say: "asks if I want a standard inspection or have I heard about the 5 star." Two things offered at once. A caller picks the cheaper one before they have heard why the other is worth it.
Fixed. The 5-Star is a recommendation, not a menu. Josh closes on the package by name. The standard inspection is what he offers after a no, never beside it.
And he must name the package in the closing question. "Want me to set you up with the 5-Star?" — not "set you up with that?". The words are how our own system knows the offer was made. An offer it cannot see is one Josh will make twice.
Adding the two towns took a minute. Looking for where to add them found something bigger.
Josh was holding two lists of Chad's towns, and they did not agree. One list had 16 towns and ended "and all points in between across the greater Houston area". The other had five: "It covers Pearland, Houston, Friendswood, League City and Galveston."
And the short list was the one he was told he could say out loud. It sits in a block that reads "these are true and you may say them", right next to "for anything else, say you do not have it in front of you". So Katy, Sugar Land, Kemah, Alvin, Angleton, Pasadena, Rosenberg, Dickinson, Seabrook, Missouri City and Lake Jackson are all towns Chad covers, and the list Josh could quote did not name one of them. Tomball and The Woodlands are the two he happened to notice.
Fixed. There is one list now, and it is the one on his form. The other place no longer names towns at all. A test stops it happening to any client.
Chad's record is already updated and live — that part needed no deploy at all.
This one is not fixed, and we are not going to pretend it is. It is his most repeated complaint.
Four different things can cut off the end of a sentence, and every one of them used to throw the words away without a word in the log. Two weeks of reports could name the symptom and never the cause. Each of the four now writes down exactly what it threw away.
The most likely answer is that it is not a fault at all, which is why we measured before fixing. All four happen when the caller talks over Josh — and people talk over a goodbye constantly. "Thanks, bye" lands on top of "have a great day". If that is what Chad heard, the right fix is to say goodbye earlier, not to stop cutting off.
What it needs: one call where Dil talks over Josh's goodbye. The log will then name which of the four it was, and the fix takes an hour.
Worth telling you, because it is the system working and not a wasted afternoon.
Sixteen helpers were set on Chad's list: one to plan each fix, and one to try to break the plan. Every plan turned out to be wrong. Several were proved wrong by running the code, not by arguing about it.
Two of them would have re-created the very fault they were fixing — by quoting the bad wording into Josh's instructions as the thing to avoid. That is how the made-up phone code got in there in the first place. One had a mistake in it that turned a test red the moment it was tried.
The lesson is the same one all week. A plan that reads well is not a plan that works. Every one of these was written by something that had read the files. The half that caught them was the half told to break them, with permission to run the code.
Ken asked whether a bigger model would make Josh better. We tried it twice on the test line. Both times every single call died — nought out of six, twice.
They failed for two different reasons, and neither was the model being too slow. The first time, the bigger model quietly switched on a "thinking" mode that the smaller one does not have, and the caller heard it as silence. The second time we turned that off — but the setting was written onto a version of Josh that could not read it. The switch was set and nothing was listening.
The only proof was one line in a log file on the server, so nobody looked. It is now on the button: press status after one call and it tells you whether the setting actually landed. That turns a wasted run of six calls into one call.
The third try is set up and costs nothing — the test line is already sitting on the right settings. Chad's and Ted's lines are untouched by all of it. That is the whole point of keeping the test line separate, and it has now been proved twice by something going badly wrong.
Three things on Chad's list still need Ken 1. Gas in the utilities question. Chad wants Josh to ask about water, electric and gas. We have not done it, on purpose. Our rule says if the utilities are off, we do not book the job. Nobody has said what to do when the water and power are on but the gas is off. Asking about gas without that answer invites Josh to turn down a job he should take. One sentence from Ken settles it.
2. Taking the card before the job, charging after. Beth checked with Stripe and it can be done. It is a build, not a setting — we have no card system wired up at all yet. It needs to go on the to-do list with a guess of how big it is.
3. Who gets a copy of the report. Ken read this one out in the meeting and moved on without ruling. Chad's setup currently says the report goes out after the payment clears. We have left it exactly as it is.
Two short jobs, and where everything lives.
Read the morning report
reports/overnight/latest.md on the overnight-reports branch.Put a night-time fix on the phone
How the next Claude session knows all this
DECISIONS-LOG.md first, every time. Today's notes are at the top: your template decision, why the e-sign was off, and what now runs on its own.HANDOFF.md. A new section 6a explains the report loop, the heartbeats and the night fixer, with the two traps we already fell into.Where things live
.github/workflows/overnight-loop.yml and scripts/overnight/README.md.austin-voice-poc/bot.py, with tests in test/slot-never-offered.test.py and test/booking-claim-guard.test.py.resources/agreements/booked-solid-standard-agreement-DRAFT.md.DECISIONS-LOG.md.api/austin/call-script.json holds the question
order and the new rules; the tests are test/utilities-known-when-occupied.test.mjs,
test/package-agreed-is-agreed.test.mjs and
test/service-area-is-one-list.test.mjs.scripts/set-service-area.mjs, which warns you what each new town costs.