We asked the phone what was wrong. It told us. Two fixes are on the phone, and the proof run passed 10 of 12.
Two things were making calls long and failing them. Both were our own software getting in Josh's way. Both are fixed in code and running on the phone.
Leak one: Josh forgot what a caller had just spelled. Leak two: the words "1800 square feet" looked like a street address to our checklist, so it stopped Josh from ever asking for the real one. Neither was Josh guessing badly. In both cases he was right, and a part we built was wrong.
This report explains each leak in plain words. It shows how we found it, what we changed, and how we know it stays fixed. At the end are the things that need a decision from Ken or Beth, and the three things I got wrong today.
Dil asked: when the test runs go bad, is it our server, or is it the voice engine? A fair question, and the answers cost different money. A bigger server is one bill. A bigger voice plan is another. So we read the server while a test run was going.
It is not the server. Memory never fell below half free. All four workers put together used about 58% of the machine. There was room.
The voice engine has a wall. Our voice plan can make sound for three sentences at the same moment. Number four waits in line. The caller hears that wait as silence. Three calls lost their voice this way, with five, ten and nine calls on the line at once.
One call proved it. A call died on a worker that had only two calls and was barely working. A quiet worker cannot make a dead mouth. A full voice account can, on every worker at once. So it was the voice engine, not the box.
But that was only three calls out of 39. I said it out loud too early. I told Dil the voice engine was why the runs looked bad. Then I read all 39 calls. The voice engine explained three failures. The other ten were ours. That is the rest of this report.
Three test suites ran this morning, one on each client line. Here is how they went.
| Line | Calls | Passed | Middle call length | Hit the 10-minute cutoff |
|---|---|---|---|---|
| GC (Chad) | 11 | 8 | 7.5 min | 0 |
| Quality (Ted) | 17 | 12 | 8.0 min | 2 |
| Staffordshire (Glen & Giselle) | 11 | 6 | 7.7 min | 3 |
Three more numbers say the same thing. Josh asked 35 questions on a middle call. The worst call had 51. A person who books inspections asks about ten. Callers said "are you still there?" 33 times. And of the 13 calls that failed, 8 failed on the same thing: getting an email or a name spelled right.
Here is a real call from this morning, on Staffordshire's line. The caller is Bernard. Watch the times.
Look at 04:41. Josh had the right answer. The caller had spelled it, and our code had handed Josh the letters. Then Josh spoke, the caller said the address again in a normal voice, and Josh went back to "Burner." Fifteen seconds after having it right.
When a caller spells a word, our code builds the letters into a note for Josh: (spelled so far: BERNARD). That part worked. But the note was wiped clean the moment Josh spoke. So on the next turn, with no letters in it, there was no note. Josh fell back on what the speech-to-text thought it heard. It thought "Burner."
The same thing hit a second call a different way. A caller said "lgrant at p s r e dot com" and then spelled "L G R A N T." Our code took the first spelled thing it found. That was "p s r e" — the part after the "at." So it told Josh the caller had spelled PSRE. Josh then guessed the name three times and gave up.
What a caller spells is now kept for the whole call. When they say the word again without spelling it, Josh's turn carries: (the caller has spelled, this call: BERNARD — use it, do not re-guess it). And when a turn has two spelled parts, the one before the "at" is the name. The one after is the website.
Both are reminders added to the caller's turn. We never rewrite what the caller said. A reminder Josh can ignore is safe. Changing the caller's words could put letters into a booking that nobody said.
We replayed Bernard's real call through the fixed code. BERNARD now stays in front of Josh on every later turn. The other call now gives LGRANT, not PSRE. Twenty-eight checks pin it, and every spelling shape that already worked still works.
This one we found by reading the phone's own log after Dil pressed the deploy button. One call, two minutes, the same line fifteen times:
Josh was trying to ask for the address. Fifteen times. Our checklist cut the question every time and asked something else instead. The caller never heard the question once. Seven calls had this shape. Hamming graded them as "repeating questions" and "never read back the address." From the caller's side, that is exactly how it looked.
Josh works from a checklist of 25 things to collect. Square footage is item five. The street address is item seven. So the caller always says the square footage first.
Our checklist has a rule for spotting an address the caller offers on their own. It looks for a number, then a street word like Road, Drive, or Square. "1800 square feet" has a number and the word "square." So the rule said: that is an address, and we have it. From then on, every time Josh asked for the real address, our gate said "you already have it" and cut the question.
We replay old calls to catch things like this. The replays never saw it, and could not. The test service writes numbers as words: "one thousand eight hundred square feet." The phone hears numbers as digits: "1800 square feet." The rule only fires on digits. So the replays read a different sentence from the one the phone heard. That is now written into our map, so nobody trusts a clean replay on a digits bug again.
The rule. A street word followed by "feet," "foot," or "footage" is never an address. "12 Rittenhouse Square" still counts. "3,200 sq ft" no longer does. Proved in both places the rule runs.
The gate. If Josh asks for the same thing and the gate cuts it twice, the gate now assumes it is the one that is wrong. It lets the third ask through. Two, not one: one cut is the gate doing its job. This is a guard for the next bad rule, whatever it turns out to be.
The prompt. The list Josh reads of "things you already have, do not ask again" no longer includes a thing the gate has just reopened. Before, the gate was letting the ask through while the prompt was forbidding it in the same breath.
Forty-eight checks. They replay a phone-style transcript with digits through the real checklist and show the address is not marked as held. They push the same ask through the gate three times and show the third one gets out. And one of them runs the old rule on purpose, to prove it really did fire on "1800 square feet."
Our rule is that test runs go three calls at a time. This morning's runs did not. One put ten calls on the line at once. The reason is a setting that looks like it caps calls but does not.
The rule to set on every run: Duration = number of cases × 160. Eleven cases is 1760. Seventeen is 2720. Both fit under the 3600 limit.
Another session measured this from the other side today, across 479 calls: runs of one to three calls at once were 6% silent. Runs of seven to nine were 26% silent. A silent call is graded as a failure. So a big run spends money and invents failures at the same time. This is the single cheapest fix on the list, and it is a setting, not code.
These are on record because the map only works if the misses are on it too.
A fault is closed only when it is fixed in code, running on the phone, and proved on a real call.
Both of today's fixes are all three. Dil pressed the deploy button twice, and we read the phone back both times to confirm the exact files landed. Then Property Masters ran twelve real test calls on the fixed phone. That is the proof run, below.
One thing this does not fix. Josh asks 35 questions because the checklist has 25 items and 19 are asked on every call. That is the design, not a bug. Which items could move to the confirmation email is a decision, below.
Property Masters was the one client line with no card run yet. So it was a clean test. Twelve calls, on the fixed phone, at 11:04 this morning. Here is the before and after. "Before" is this morning's 39 calls on the other three lines.
| What we measured | Before (39 calls) | Proof run (12 calls) | Verdict |
|---|---|---|---|
| Calls that passed | 55% to 73% | 83% (10 of 12) | better |
| Address question cut by our gate | 15 in a row on one call | 0 on the whole run | closed |
| Failures that name email spelling | 8 of 13 | 0 of 2 as the main cause | closed |
| Calls that hit the 10-minute cutoff | 5 of 39 | 1 of 12 | better |
| "Are you still there?" per call | 0.85 | 0.50 | better |
| Questions per call (middle) | 35 | 35 | same, as expected |
| Silence, share of the call | 41% | 42% | same, as expected |
| Middle call length | 7.5 min | 7.6 min | same, as expected |
| Calls on the line at once | 5 to 6 | 5.4 | still over the wall |
The two leaks are closed. The address question was cut zero times. This morning it was cut fifteen times on one call. Of the two calls that failed, neither failed on a caller spelling an email five times. The new gate guard fired five times, always on the second cut, and let the next question through every time. It did its job on names, not addresses, which is exactly what a guard for the next bad rule is for.
The three numbers that did not move, did not move. Questions, silence, and call length are the same. I said they would be. They come from the checklist having 25 items, not from either leak. That is Ken's decision, above.
One thing was not done. This run still put five or six calls on the line at once. The Duration setting was not changed. So the voice wall was still there. It passed 10 of 12 anyway. With Duration set to cases × 160, the next run should do better still.
One caller said "Yes, I am working with a buyer's agent." Our checklist decided she had no agent and cut "what's your agent's name?" eight times on one call. Josh got the name in the end, a minute and a half late. Same loop shape as the address, different rule.
Callers say "three zero zero two two." Our checklist did not file that as a ZIP on two calls, and asked again. On one call Josh read the ZIP back as "thirty thousand and twenty-two." A ZIP is five digits, never a number.
Josh tried to offer the Peace of Mind package again after the caller had already heard it. Our gate cut it every time, which is right. But he kept trying. Watching whether that costs turns.
None of these is fixed yet. They are here because the map only works if the open leaks are on it. The first two are the same shape as today's: Josh was right, and a rule of ours filed the answer wrong.
Standing rule, in every report: we do not text homeowners. The only text we send is the booking alert to the inspector's own phone.
"We read the phone while a test run was going. The server is fine. The voice engine has a three-call wall, and that cost us three calls out of 39. The other ten failures were our own software, and both causes are fixed."
"Josh was right both times. He had the spelled name and we threw it away. He tried to ask for the address and we cut him off fifteen times. The fixes are in code, on the phone tonight, and the next run on Property Masters is the proof."
"The cheapest fix today was a setting. Our test runs were putting ten calls on a line that can voice three. Set every run to cases times 160 and the silent failures go away for free."