Daily report · Friday 4 September 2026

Nine faults found in Josh. Six fixed in code, not in a prompt. And an answer on what to buy.

You asked, in the meeting, whether the fixes hold — and what we are doing here if the same problems keep coming back. Today we stopped guessing and started measuring: eighteen test calls instead of one. They found nine faults. Six are fixed in code and on the phone. The most useful thing we learned is why the same problems keep coming back, and it is not what any of us assumed. Chad's and Ted's lines were never touched all day.

18test calls, not one
9faults found
6fixed in code today
124tests that never ran, now do
1decision that needs you

Read this first No customer of Chad's or Ted's spoke to Josh today, and nothing we changed reached their lines without being tested first. Every experiment ran on the 888 bench. We proved that the hard way: six test calls collapsed completely on the bench this afternoon and both client lines carried on working, untouched. That separation has existed for weeks. Today was the first time something actually went wrong while it was holding.

The short version


Your question in the meeting was the right one. You said "do those fixes hold?" and "what he thinks are easy fixes still generate a list that Chad just gave us." We could not answer that honestly, because we had been fixing things and hoping. So today we ran eighteen synthetic callers through Josh instead of making one phone call and forming an impression.

The answer is that some fixes held and some did not — and now we know which. The Friday rule holds. The area code holds. The mould refusal holds. The email question that Chad's caller never got asked is fixed and proven. But the spelled-out email still comes back wrong, and the sign-off still cuts off — "have a great…" — exactly as Chad reported.

And here is the part worth your time. We counted where today's nine faults actually came from. Six of them were our own code — not the AI being stupid, not the phone line, not the model. A missing word in one line that hung the booking. A guard we built that was deleting Josh's own questions. That is the honest answer to "we can build space stations but we cannot fix Josh."

What is fixed and on the phone right now

All six are code, not prompt lines — your ruling from the meeting. Each has a test that goes red if it comes back.


1

The booking no longer hangs at the last moment

Chad's #15: "drug a bit on the scheduling of the inspection, said he couldn't schedule it on his end." We found it. One line was missing a single word, so when Josh tried to save the booking the system waited for an answer that never came. The caller heard "still with you" three times and then "I can't finish the booking on my end." They had given every detail.

Verified on a live call. The same moment now completes and Josh carries on talking.

On the phoneProven on a real call
2

Josh's own questions were being deleted before the caller heard them

This one explains four separate items on Chad's list at once: "fails to get email address", "never asked for my name", the "I'm still here" loops, and the zip code you raised.

We had built a guard to stop Josh asking the same question twice. It worked — too well. If a sentence asked for two things and Josh already had one of them, the guard threw away the whole sentence, including the part we still needed. On one of Chad's calls Josh tried to get an email four times and the caller heard none of the four.

On the phone4 punch-list items
3

A phone number the system rejects now gets asked for again

Found on a live test call. A caller gave a number that could not exist. The system correctly refused it and told Josh to ask again — and the guard above then binned every one of his six attempts as a repeat. The caller said "you already have the number, can you look it up?" and hung up.

Fixed and then proven by a synthetic caller doing the same thing, which is exactly what these tests are for.

On the phone
4

Josh has the checklist you asked for

Your words in the meeting: "maybe what we do is Josh uses a checklist… we start working down more in an organized way, like the ISN, and they can ask questions as we go along, and he can come back to it."

Built. Seven things a booking cannot exist without, nine more the office wants, and three that only apply sometimes — the build phase on new construction, for instance. It is arithmetic in the code, not a reminder in the prompt, so Josh cannot forget it. He called it unprompted on a live call and got back a clean answer.

⚠️ One deliberate choice: a caller who refuses an email still gets booked. A refusal must never cost the job — the office chases the blank.

On the phoneYour ask, from the meeting
5

Josh told a home buyer we would text him. Now he cannot.

A caller refused to give an email, so Josh said the agreement and payment link would come by text. We do not text homeowners — our number is registered with the carrier for inspector alerts only, and our own website publishes the promise.

Nothing was actually sent, so there is no fine. But a caller was told, on a recorded call, that we would do the one thing we publicly promise never to do.

The rule was already in the prompt, in capitals. It did not hold, because the caller had refused an email and Josh needed somewhere to send the agreement. That is your ruling proving itself: "a prompt line is only a strong suggestion." It is a code guard now.

On the phoneCarrier risk closed
6

The area code Josh kept inventing was in his own instructions

Six calls across both clients: a caller said 281 and Josh read back 713. Five of those were on Ted's Delaware line, where nobody said 713 at all — it is a Houston area code.

It comes from us. 713-555-0170 appears nine times in Josh's own instructions — eight of them inside the worked example that teaches him not to mangle a read-back. The lesson was the source.

We left the lesson alone (it fixed a real fault, and rewriting it risks re-opening that one) and added a guard instead: if the last seven digits are right and the first three are wrong, Josh stops and asks again rather than asserting it.

On the phoneBoth clients affected

The number that answers your question about the phone line

You asked: "is it the telephone sucks, and can we move up the quality on that end of the stack?" Partly — and now we know by how much.


Three test runs, same code, same hour, same day. The only thing that changed was how many callers were on the line at once.

Callers at onceSlowest reply"One moment, I'm still here"Calls that finished
35.5 seconds0all of them
66.1 seconds0all of them
914.0 seconds16one caller hung up

This is not a testing problem. It is a launch problem. Nine callers at once are nine callers at once, whether a test dialled them or Texas did. One of those calls was dead before the caller had finished saying her name — she heard silence, then "I'm still here", then nothing, and hung up. That is what a real customer would hear.

Chad and Ted today are fine. Ten clients are not. This is the clearest thing we learned all day and it belongs on the launch list.

The cause is ours and it is fixable with money: Josh runs four workers, and each one can carry roughly one call before the next has to queue. It also explains the "I'm still here" loops you have been reading about in Chad's notes for weeks.

Ken — the upgrade question, answered straight

You asked twice whether a higher-level model would do it better or faster, and said you would spend the money. Here is what the evidence says.


We tested it. Twice. Both times all six calls died.

What Josh runs nowThe bigger model
Time before he speaks2.3 – 3.9 sec4.5 – 11.4 sec
Typical replyabout 2.5 secup to 12.7 sec
Calls completed6 of 60 of 6

But the failure was not the model being too slow — it was a setting we did not know had changed. The bigger model, unlike the current one, stops and thinks before every sentence unless you tell it not to. On a phone call the caller hears that as silence. We have since built the switch that turns it off, and the test can be re-run properly.

The honest recommendation: do not buy a bigger brain yet We counted where today's nine faults came from. Six were our own code — a missing word, a guard eating questions, a rejected value that never reset, dead air from server load, a truncated sign-off, a missing pattern. A better model fixes none of those. Two were maybes.

So the answer to "we can build space stations but we cannot fix Josh" is not that the AI is not good enough. It is that we were repairing a machine by rewriting its instruction manual — which is exactly what your own code-versus-prompt ruling says to stop doing. Today was the first day we followed it properly.

What we would spend money on, in order:

Where the money actually goes

  1. More server capacity — the one thing that is measured. Four workers, nine callers, calls dying. This is the only place today's evidence says money buys a guaranteed result.
  2. Two speed settings we have never once tried — free. Both sit on a button we already have. One of them stops Josh's voice being re-processed on every single reply. Nobody has ever measured either.
  3. The voice itself. We are on the expensive, slow voice model — deliberately, after the cheap one sounded worse in August. But that test was about hearing, not speaking. There is a faster voice we have never actually tried, and it is half the price.
  4. A bigger brain — later, and only if the numbers say so. The re-test is built and ready. If it comes back level, the money belongs in items 1 to 3.

One thing to clear up: "Ultracode" is not a Josh upgrade You spotted it in Claude Code and asked whether it would help. It is a setting that changes how the development assistant works — it makes building and checking fixes more thorough. It runs nothing on a phone call. It can make us faster at fixing Josh; it cannot make Josh better while a customer is talking to him.

What is still broken, in Chad's own words

Honest list. These are on his sheet and they are not fixed yet.


Chad saidWhere it stands
"have a great…" then nothing · "Thanks home inspection"⚠️ Still happening. The sign-off cuts off on nearly every call. Next on the list
Garbled spelling, "the gibberish is back"⚠️ Still. A caller spelled an email in code words and Josh read the code words back as if they were the address
Utilities say water and electric, not gasConfirmed — one-line fix, in hand
Asks about utilities when the home is occupiedNot done. If someone lives there, they are on
"Is this for you as the buyer or for something else"Not done. It is someone else, as you said
Repeats the 5-star after I already agreedNot done
Should lead with the 5-star, standard only as a down-sellCause found — see below
Tomball and The Woodlands must be coveredConfirmed missing from Chad's record. Going in
Card authorised before, charged 2 hours after the startBeth confirmed with Stripe it is possible. It is a build, not a setting

The 5-star finding, because it is not what anyone assumed The either/or is not in Chad's setup. His record already says "Want me to set you up with the 5-Star?" — closing on the package alone. We checked all seventeen inspector records; not one contains a choice between the package and the standard.

Josh is adding it himself. Four places in his instructions tell him what order to say things in, and not one tells him the package is a recommendation rather than a menu. There is a recorded call that proves it: the first two sentences are Chad's script word for word, and the "or would you prefer the standard inspection" is Josh's own addition.

Rewriting Chad's script would never have fixed this.

Dil is on Josh this weekend


Dil's focus for the weekend is Josh and nothing else. There is a practical reason as well as an obvious one.

Every time anyone pushes work to the system, the main application restarts for a second or two. That is normal and it is how work goes live. But Josh depends on that application during a live call — it is where he gets the company he is answering for, the prices, the diary and the booking.

Ninety of those restarts happened today, from several people working at once. Four of them landed inside an eleven-minute test run. That is not the whole story of today's failures — we traced each one to its own cause and can show the work — but it is real interference, and it makes a bad result harder to read.

What that means in practice Josh's testing needs a quiet system. This weekend Dil runs the Josh work with the other streams paused, so a number that comes back bad is bad because of Josh and not because somebody shipped a report while a caller was mid-sentence.

The booking itself is already protected — it retries over a restart, and has since August. The weaker spot is the very start of a call, where Josh fetches the company he is answering for. Hardening that is on the list.

What we are testing next


In this order

  1. The speed setting nobody has tried — one button, no cost, and it addresses the delay directly.
  2. The bigger model, properly this time — with the thinking switched off. The button now refuses to run unless the right code is on the server, so today's mistake cannot repeat.
  3. The sign-off and the spelled-out email — the two loudest things still on Chad's sheet.
  4. Then a full re-run of both test suites, one at a time, six cases or fewer — which is the limit we now know the server can take.

The one decision that needs you Should Josh lead with the 5-star every time and mention the standard inspection only as a down-sell? Dil has already said yes on your behalf and we are building it. Say the word if you want it any other way — it is a sales decision, not a technical one, and it changes what every caller hears.

Something we got wrong today


Worth writing down, because it is the same failure as the ones above.

Our test suite had been stopping a third of the way through, for days. It runs 217 checks in a row, and one failing check in the middle stopped everything after it — so 124 tests never ran at all and "the tests pass" did not mean what it sounded like. Found by the other team stream. Fixed: every check now runs and every failure is reported.

The moment it was fixed, it caught a mistake of ours from earlier the same day — a fix that had quietly re-created the exact fault it was written to repair. That is the argument for measuring rather than assuming, in one sentence.