Daily report · Friday 4 September 2026
You asked, in the meeting, whether the fixes hold — and what we are doing here if the same problems keep coming back. Today we stopped guessing and started measuring: eighteen test calls instead of one. They found nine faults. Six are fixed in code and on the phone. The most useful thing we learned is why the same problems keep coming back, and it is not what any of us assumed. Chad's and Ted's lines were never touched all day.
Read this first No customer of Chad's or Ted's spoke to Josh today, and nothing we changed reached their lines without being tested first. Every experiment ran on the 888 bench. We proved that the hard way: six test calls collapsed completely on the bench this afternoon and both client lines carried on working, untouched. That separation has existed for weeks. Today was the first time something actually went wrong while it was holding.
Your question in the meeting was the right one. You said "do those fixes hold?" and "what he thinks are easy fixes still generate a list that Chad just gave us." We could not answer that honestly, because we had been fixing things and hoping. So today we ran eighteen synthetic callers through Josh instead of making one phone call and forming an impression.
The answer is that some fixes held and some did not — and now we know which. The Friday rule holds. The area code holds. The mould refusal holds. The email question that Chad's caller never got asked is fixed and proven. But the spelled-out email still comes back wrong, and the sign-off still cuts off — "have a great…" — exactly as Chad reported.
And here is the part worth your time. We counted where today's nine faults actually came from. Six of them were our own code — not the AI being stupid, not the phone line, not the model. A missing word in one line that hung the booking. A guard we built that was deleting Josh's own questions. That is the honest answer to "we can build space stations but we cannot fix Josh."
All six are code, not prompt lines — your ruling from the meeting. Each has a test that goes red if it comes back.
Chad's #15: "drug a bit on the scheduling of the inspection, said he couldn't schedule it on his end." We found it. One line was missing a single word, so when Josh tried to save the booking the system waited for an answer that never came. The caller heard "still with you" three times and then "I can't finish the booking on my end." They had given every detail.
Verified on a live call. The same moment now completes and Josh carries on talking.
This one explains four separate items on Chad's list at once: "fails to get email address", "never asked for my name", the "I'm still here" loops, and the zip code you raised.
We had built a guard to stop Josh asking the same question twice. It worked — too well. If a sentence asked for two things and Josh already had one of them, the guard threw away the whole sentence, including the part we still needed. On one of Chad's calls Josh tried to get an email four times and the caller heard none of the four.
Found on a live test call. A caller gave a number that could not exist. The system correctly refused it and told Josh to ask again — and the guard above then binned every one of his six attempts as a repeat. The caller said "you already have the number, can you look it up?" and hung up.
Fixed and then proven by a synthetic caller doing the same thing, which is exactly what these tests are for.
Your words in the meeting: "maybe what we do is Josh uses a checklist… we start working down more in an organized way, like the ISN, and they can ask questions as we go along, and he can come back to it."
Built. Seven things a booking cannot exist without, nine more the office wants, and three that only apply sometimes — the build phase on new construction, for instance. It is arithmetic in the code, not a reminder in the prompt, so Josh cannot forget it. He called it unprompted on a live call and got back a clean answer.
⚠️ One deliberate choice: a caller who refuses an email still gets booked. A refusal must never cost the job — the office chases the blank.
A caller refused to give an email, so Josh said the agreement and payment link would come by text. We do not text homeowners — our number is registered with the carrier for inspector alerts only, and our own website publishes the promise.
Nothing was actually sent, so there is no fine. But a caller was told, on a recorded call, that we would do the one thing we publicly promise never to do.
The rule was already in the prompt, in capitals. It did not hold, because the caller had refused an email and Josh needed somewhere to send the agreement. That is your ruling proving itself: "a prompt line is only a strong suggestion." It is a code guard now.
Six calls across both clients: a caller said 281 and Josh read back 713. Five of those were on Ted's Delaware line, where nobody said 713 at all — it is a Houston area code.
It comes from us. 713-555-0170 appears nine times in Josh's own instructions — eight of them inside the worked example that teaches him not to mangle a read-back. The lesson was the source.
We left the lesson alone (it fixed a real fault, and rewriting it risks re-opening that one) and added a guard instead: if the last seven digits are right and the first three are wrong, Josh stops and asks again rather than asserting it.
You asked: "is it the telephone sucks, and can we move up the quality on that end of the stack?" Partly — and now we know by how much.
Three test runs, same code, same hour, same day. The only thing that changed was how many callers were on the line at once.
| Callers at once | Slowest reply | "One moment, I'm still here" | Calls that finished |
|---|---|---|---|
| 3 | 5.5 seconds | 0 | all of them |
| 6 | 6.1 seconds | 0 | all of them |
| 9 | 14.0 seconds | 16 | one caller hung up |
This is not a testing problem. It is a launch problem. Nine callers at once are nine callers at once, whether a test dialled them or Texas did. One of those calls was dead before the caller had finished saying her name — she heard silence, then "I'm still here", then nothing, and hung up. That is what a real customer would hear.
Chad and Ted today are fine. Ten clients are not. This is the clearest thing we learned all day and it belongs on the launch list.
The cause is ours and it is fixable with money: Josh runs four workers, and each one can carry roughly one call before the next has to queue. It also explains the "I'm still here" loops you have been reading about in Chad's notes for weeks.
You asked twice whether a higher-level model would do it better or faster, and said you would spend the money. Here is what the evidence says.
We tested it. Twice. Both times all six calls died.
| What Josh runs now | The bigger model | |
|---|---|---|
| Time before he speaks | 2.3 – 3.9 sec | 4.5 – 11.4 sec |
| Typical reply | about 2.5 sec | up to 12.7 sec |
| Calls completed | 6 of 6 | 0 of 6 |
But the failure was not the model being too slow — it was a setting we did not know had changed. The bigger model, unlike the current one, stops and thinks before every sentence unless you tell it not to. On a phone call the caller hears that as silence. We have since built the switch that turns it off, and the test can be re-run properly.
The honest recommendation: do not buy a bigger brain yet We counted where today's nine faults came from. Six were our own code — a missing word, a guard eating questions, a rejected value that never reset, dead air from server load, a truncated sign-off, a missing pattern. A better model fixes none of those. Two were maybes.
So the answer to "we can build space stations but we cannot fix Josh" is not that the AI is not good enough. It is that we were repairing a machine by rewriting its instruction manual — which is exactly what your own code-versus-prompt ruling says to stop doing. Today was the first day we followed it properly.
What we would spend money on, in order:
Where the money actually goes
One thing to clear up: "Ultracode" is not a Josh upgrade You spotted it in Claude Code and asked whether it would help. It is a setting that changes how the development assistant works — it makes building and checking fixes more thorough. It runs nothing on a phone call. It can make us faster at fixing Josh; it cannot make Josh better while a customer is talking to him.
Honest list. These are on his sheet and they are not fixed yet.
| Chad said | Where it stands |
|---|---|
| "have a great…" then nothing · "Thanks home inspection" | ⚠️ Still happening. The sign-off cuts off on nearly every call. Next on the list |
| Garbled spelling, "the gibberish is back" | ⚠️ Still. A caller spelled an email in code words and Josh read the code words back as if they were the address |
| Utilities say water and electric, not gas | Confirmed — one-line fix, in hand |
| Asks about utilities when the home is occupied | Not done. If someone lives there, they are on |
| "Is this for you as the buyer or for something else" | Not done. It is someone else, as you said |
| Repeats the 5-star after I already agreed | Not done |
| Should lead with the 5-star, standard only as a down-sell | Cause found — see below |
| Tomball and The Woodlands must be covered | Confirmed missing from Chad's record. Going in |
| Card authorised before, charged 2 hours after the start | Beth confirmed with Stripe it is possible. It is a build, not a setting |
The 5-star finding, because it is not what anyone assumed The either/or is not in Chad's setup. His record already says "Want me to set you up with the 5-Star?" — closing on the package alone. We checked all seventeen inspector records; not one contains a choice between the package and the standard.
Josh is adding it himself. Four places in his instructions tell him what order to say things in, and not one tells him the package is a recommendation rather than a menu. There is a recorded call that proves it: the first two sentences are Chad's script word for word, and the "or would you prefer the standard inspection" is Josh's own addition.
Rewriting Chad's script would never have fixed this.
Dil's focus for the weekend is Josh and nothing else. There is a practical reason as well as an obvious one.
Every time anyone pushes work to the system, the main application restarts for a second or two. That is normal and it is how work goes live. But Josh depends on that application during a live call — it is where he gets the company he is answering for, the prices, the diary and the booking.
Ninety of those restarts happened today, from several people working at once. Four of them landed inside an eleven-minute test run. That is not the whole story of today's failures — we traced each one to its own cause and can show the work — but it is real interference, and it makes a bad result harder to read.
What that means in practice Josh's testing needs a quiet system. This weekend Dil runs the Josh work with the other streams paused, so a number that comes back bad is bad because of Josh and not because somebody shipped a report while a caller was mid-sentence.
The booking itself is already protected — it retries over a restart, and has since August. The weaker spot is the very start of a call, where Josh fetches the company he is answering for. Hardening that is on the list.
In this order
The one decision that needs you Should Josh lead with the 5-star every time and mention the standard inspection only as a down-sell? Dil has already said yes on your behalf and we are building it. Say the word if you want it any other way — it is a sales decision, not a technical one, and it changes what every caller hears.
Worth writing down, because it is the same failure as the ones above.
Our test suite had been stopping a third of the way through, for days. It runs 217 checks in a row, and one failing check in the middle stopped everything after it — so 124 tests never ran at all and "the tests pass" did not mean what it sounded like. Found by the other team stream. Fixed: every check now runs and every failure is reported.
The moment it was fixed, it caught a mistake of ours from earlier the same day — a fix that had quietly re-created the exact fault it was written to repair. That is the argument for measuring rather than assuming, in one sentence.