Daily report · Sunday 7 September 2026
Two jobs ran today, side by side. On Chad’s real line we found six faults, fixed them, and proved it by running the same nine calls again — his score went from 70 to 81. On the test line, a second Josh started the day finishing nothing and ended it booking five calls out of six. Nothing was sent to a client, and no customer heard any of it.
Read this first We do not text homeowners. The only text this system sends is the booking alert to the inspector's own mobile. A home buyer, seller or agent hears from us by email only. That did not change today and it is not going to.
Nothing reached a client. Every test call went to Chad's line on purpose, which is what Ken asked for. Test bookings landing in the real system are expected.
It helps to know there are two Joshes. They are both in this report.
| Josh on Chad’s line | Josh on the test line | |
|---|---|---|
| Who calls it | Chad’s real customers | Nobody. It is a test number and a made-up company |
| How he works | He decides what to ask next | He follows a checklist, like Ken asked for |
| Who says the numbers | Josh says them | The software says them. Josh does the talking |
| Today | Six faults fixed and live | Six test runs, from stuck to booking |
The test line is where we try things first. If a change works there, it moves to Chad’s and Ted’s lines later, one piece at a time. Nothing moves over until it has earned it.
This is the phone real customers call. Everything below is live on it now.
We ran the same nine calls twice. Same script, same background noise, same everything. The only thing that changed in between was the code.
| What we measured | Before | After |
|---|---|---|
| Overall score | 70 | 81 |
| Josh does not waste the caller's time | 2 of 9 | 6 of 9 |
| Josh sounds like he is listening | 0 of 9 | 4 of 9 |
| Josh asks again for an email he already has | 1 call | 0 calls |
| Price matches the price tool | 9 of 9 | 9 of 9 |
| Never guesses at what the inspection will find | 9 of 9 | 9 of 9 |
| Booking finished cleanly | 7 of 9 | 6 of 9 |
One row went the wrong way, and we chased it "Booking finished cleanly" went from 7 to 6. Two calls got longer instead of shorter. We checked whether our own fixes caused that before doing anything else. They did not — calls were 4 seconds longer on average, which is noise. What we found instead was a real fault, and it is item 5 below.
All six are on Chad's phone right now. Each one came from a real call we can point at.
A caller spelled her email as "r, dot, t for tango, i for India…" and Josh wrote down
RTILGHMAN. The dot was thrown away. He then had to guess where it went, and he
could not.
This is Chad's "the gibberish is back" complaint, and it had been blamed on a bad phone line since 27 August. It was not the line. It was one thrown-away character. Dots now survive in 6 of 7 emails, up from none.
Ten calls out of a hundred had the caller say the street address in their first sentence. Josh did not hear it as an address, so he asked again. Callers find this maddening, and it is the single biggest reason his "does not waste your time" score was 2 out of 9.
He hears it now. That score went from 2 out of 9 to 6 out of 9.
This is the worst one, and it is worth reading in full. It took four tries and two and a half minutes to get one email right. The caller finally confirmed it:
Josh: "…r for Romeo, t for tango, i for India… is that correct?"
Caller: "Yes, that's correct. Thank you for double-checking."
Josh, five seconds later: "You're welcome. And what's your email address?"
He had it. She confirmed it. He asked again from scratch, got it wrong, and gave up with "the team will reach out." Two and a half minutes of her time, thrown away by a yes.
Fixed. On the repeat run this happened zero times.
We found this while fixing item 3, and it is the kind of fault that hides for months. Our code looks for phrases like "no, it's…" to know the caller is correcting Josh. But the apostrophe in the phone system's text is a curly one, and our code was looking for a straight one. They are different characters.
So three out of six real corrections were invisible. Josh carried on with the wrong answer and nothing told him. Fixed in the one place all our word-matching runs through, so it cannot come back for a different word list.
This is what made that call run to ten minutes. The caller said "Of course. Thank you." and waited. Josh owed her the street address. Instead he said:
"One moment — I'm still here." (07:49)
"One moment — I'm still here." (08:07)
"One moment — I'm still here." (08:25)
He asked the real question at 09:07 — ninety seconds later. By then the appointment slot he had offered her was gone, and his next words were "that 1:00 on Friday went while we were talking."
The cause: we have two safety nets for silence, and only one of them knew how to ask a question. The one that fires when the caller is waiting was the one that could not. Both ask now.
We have a rule that stops Josh asking the same thing over and over. It is a good rule — one caller was once asked for her ZIP code four times after answering it correctly the first time.
But the rule counted how often Josh asked, not how often the caller answered. So a caller who said "Of course. Thank you." used up one of Josh's two goes without answering anything. Two polite replies and Josh went silent for good.
The rule now counts answers. A "thank you" or an "I don't know" no longer burns a turn. We tested this in both directions: the old four-times-ZIP fault cannot come back either.
A second Josh runs on a test number, for a made-up company called XYZ Home Inspections. Nobody real ever calls it. This is where the checklist Ken asked for is being built.
He works down a list of questions, one at a time, in order. He still talks like Josh and still answers questions. But the software says every number, every price and every appointment time, and it only says “you’re all set” after the booking has really saved.
We ran him six times today. Each time we read what went wrong, fixed it, and ran him again. Here is the first run of the day next to the last one.
| On the test line | This morning | This afternoon |
|---|---|---|
| Score | 22 | 67 |
| Sound and connection score | 21 | 81 |
| Calls that got the job done | 0 of 3 | 6 of 6 |
| Calls where nothing was made up | failed | 6 of 6 |
| Calls that booked and closed properly | 0 | 5 of 6 |
| Times Josh spoke in one call | about 120 | 41 to 52 |
| Longest the caller ever waited | 129 seconds | under 6 seconds |
The two rows we care about most “Calls that got the job done” and “nothing was made up” both passed on all six calls. The price was right every single time. That is the whole reason the test line exists: the software holds the pen for anything with a number in it, so a wrong price is not caught after the fact — it never gets said.
Six fixes, all from real test calls we can point at. They went out through the day, and each run showed whether the last fix worked.
This was the worst one, and it killed two whole calls. Josh asked “is this a standard home inspection?” about forty times over four minutes. The caller answered twice. He carried on asking.
The cause was a waiting rule deep in the software we use. It would not move to the next question until everything Josh had said was finished playing. With noise in the background he was always saying something, so it never moved.
Now the checklist moves the moment an answer comes in. On the afternoon runs the calls went from start to finish with no sticking at all.
When Josh did not hear an answer, he asked again straight away. Then again. Each new question got in line behind the last one, so he was still asking while the caller was answering.
Now he waits ten seconds, asks three times at most, then stops and listens. If the background is noisy he says so once, in his own words, instead of just asking again.
One test had a hospital conversation playing behind the caller. Someone in it said a blood pressure reading, a co-pay and an appointment time. Josh took ten digits from three different sentences, stuck them together, and read them back as the caller’s phone number.
Now a phone number has to be said in one go, the way a real person says it. Numbers scattered across a few sentences are not a phone number any more.
When Josh cannot catch something he sometimes writes a stand-in word like Unknown. The software then read it out: “Unknown, unknown, seven five seven three — did I get that right?”
Now a stand-in word is never read back. Josh asks the question again instead.
Four times in three calls Josh said “I think you may have reached the wrong number.” Every one of those callers had dialled exactly right. He was reacting to the chatter in the background.
Now he can only say that if the caller says it first.
Five of the six good calls ended with “You’re all set. The office will send the details by email.” Ken ruled that out on 6 September: most inspectors have no office, and more than half work alone. We had taken it out of one line and missed the other.
Now Josh says only what really happens. If we have an email, a confirmation really is on its way. If we do not, the inspector really does get the booking on their own phone. Nothing else is promised. The exact words are Ken’s to choose — what is in there now is a stand-in.
One more thing we caught, and it is the kind that hides for months On one call the caller corrected her last name, and Josh repeated the whole booking back to her himself — including the zip code, read out as “seventy seven thousand five hundred eighty four” instead of “seven, seven, five, eight, four.”
The software is supposed to be the only thing that says numbers. Two guards should have stopped him. Neither did: one was looking for digits and he used words, and the other was looking for phrases like “you’re all set” and he said “so that’s an inspection for…” instead. Both are closed now. This is the same kind of hole that lets a wrong price through, so it was worth stopping today rather than next month.
Four faults are open, two on each line. We would rather write them down than leave them for someone to find on a real call.
This is the root of the email problem, and today's work did not fix it. It made Josh recover faster from it.
| Caller said | Josh heard |
|---|---|
| Tilghman | Todman, Toddman, Tillman |
| Whitlock | Wheatlock, witlock, wheatlot |
| Kerr | kur, care, curse |
| Trowbridge | Traubridge |
There is an easy-looking fix we are not taking: we could add these exact names to the system's word list. They are test names, so that would make the test pass without making a real call any better. That is cheating our own exam.
Asking the caller for "R for Romeo" works. When Josh finally did it, he got the hard email right first time. Today we made him ask on the second try instead of the third.
But it only fires on questions our code writes for him, not on ones he writes himself. On one call he was corrected three times on the same name, never asked for letter-words, and gave up. That is the next job.
Benchley came out as Finchley and Century on three of six calls. Every other answer gets read back to the caller to check. The name does not. That is the next real job on the test line.
Every test call failed on one thing: length. The checklist has 23 questions, and asking 23 questions takes about six minutes. That is not a bug. It is a fair question for Ken: which of those 23 does a first phone call really need? The rest could be asked later, by email or in the portal.
Written down because they cost real time, and because a report that only lists wins is not worth reading.
We blamed a broken test run on the wrong thing Eleven test calls died with Josh cut off mid-sentence. We blamed a server restart that had happened at the same moment. Dil re-ran it with no restart and it failed exactly the same way. The real cause was the background noise being set too high — Josh was answering the traffic. Dil had said the noise sounded too loud when he listened to it. He was right and we should have started there.
We said a money fault was open when it was already fixed Josh quoted $425 for a 2,200 sq ft home when Chad's own table says $450, on two real calls, and one of them booked. We reported it as an open bug twice. It was fixed on 6 September — but the phone had not been updated since 4 September, so the fix was sitting on the shelf. Pressing the deploy button today is what shipped it. This is exactly why we built a guard for that gap today.
We said Ted's line was not being tested. It was. The first draft of this report told Ken that Quality's line had no automatic testing and asked him to approve adding it. It has been on the same schedule as Chad's all along. We had looked at one screen, seen one line on it, and not checked. Dil caught it before this went out. The real finding is the opposite and it matters more: both lines are already running, and that costs about $1,685 a month against a $1,200 budget.
We nearly shipped a fix that made things worse The first version of fix 3 was too broad. Tested against 112 recorded calls it started silencing Josh almost twice as often, and it ate a safety line about utilities being switched off. We caught it because our test suite runs every check instead of stopping at the first failure. The narrower version is what shipped.
Three decisions
The evidence
What happens tonight, with nobody watching Both lines get six test calls at 6 p.m. Texas time — Chad's and Ted's. Our checker reads them, writes a report, and drafts fixes overnight. That is the first real-world test of the $425 price fix. The report will be waiting in the morning.