Investigation · started Thursday 20 August 2026
Dil asked for the call records first, and he was right to. Six hundred and twenty-two calls are on file. Isolating the tester’s gives the sessions, the times and the durations — and it kills the leading suspect for the cut-outs. The releases are not doing it. What is left is two candidates, one command that separates them, and a switch built to settle exactly this that has never been thrown.
The tester is one number, and it has rung Josh 47 times. Three sessions: Saturday 15 August, Tuesday 18 August and Wednesday 19 August. Twenty-three of the 47 booked.
The release theory cannot carry it — but it is not zero, and my first pass
said it was. Correcting that here: a first count compared timestamps as text, and six of
the commits carry a +08:00 offset. Redone on real instants, two pushes landed
inside a live call. Both are below, with the calls they landed in.
It still cannot be the main cause. Two calls, against eleven cut-outs in the notes — and the Saturday session, 22 of the 37 calls, had not one commit to main in a twenty-four hour window either side.
Two suspects are left, and they are both in the phone bot. A false barge-in — Josh hearing something and yielding mid-sentence — or the number filter mangling his speech on the way out. One command tells them apart, and the answer for nine of those calls is sitting in the server’s log right now.
One caller, three evenings. Times are Central, the way GC works.
| Session | Calls | Booked | Window | Median length | Longest |
|---|---|---|---|---|---|
| Fri 14 Aug | 1 | 0 | 17:58 | 8m 07s | 8m 07s |
| Sat 15 Aug | 22 | 10 | 13:45 – 19:08 | 5m 04s | 9m 13s |
| Tue 18 Aug | 15 | 9 | 21:18 – 22:55 | 4m 28s | 6m 12s |
| Wed 19 Aug | 9 | 4 | 06:25 – 23:09 | 3m 43s | 8m 27s |
The thirty-six notes are almost certainly the 15th and the 18th. Those two sessions are 37 calls, and one of them lasted twelve seconds — which is a redial, not a scenario. This needs one line back from the tester rather than a guess, because the whole analysis below hangs on which nights are being described. If the notes are the 18th and 19th instead, one of the two suspects changes.
The tester books 49% of the time. Real callers on the same line book 58%. That gap is expected — several scenarios are written to hang up once a price is given — but it is the baseline any future run should be read against.
Four calls ended in under thirty seconds — 10, 10, 12 and 14 seconds. Call 22 in the notes is the one that reads “I never finished the call because I was so frustrated”, and a call that short is somebody hanging up, not a conversation that failed.
Dil’s question, and it has three separate answers because there are three separate kinds of restart.
“A restart” covers three different things here, and only one of them touches a call. Separating them is most of the answer.
| Kind of restart | What it restarts | Did it hit these calls? |
|---|---|---|
| A release — every push to main | The office app only. pm2 restart q. It does not touch the phone bot | Yes, twice — both inside a live call |
| A phone deploy — copied by hand | The phone bot itself. Kills any call in progress | No. Not one during any session |
| The machine running out of memory | Whatever it picks. Can take everything | No. All of them missed the sessions |
The first pass on this said zero, and it was wrong. It compared timestamps as text, and
six of the commits are written with a +08:00 offset, so they sorted into the wrong
place. Redone properly:
| Session | Calls | Pushes during it | Landed in a live call? |
|---|---|---|---|
| Sat 15 Aug | 22 | 0 — and none for 24h either side | — |
| Tue 18 Aug | 15 | 1, at 03:16:44Z | Yes — 24 seconds before a call ended |
| Wed 19 Aug | 9 | 1, at 03:33:33Z | Yes — about two thirds through a call |
What a release actually does to a call in progress. It does not restart the phone bot, so Josh keeps talking. It restarts the app he fetches the price list, the diary, the caller lookup and the booking from. For a second or two those are a closed door.
Every one of those lookups was a single attempt until then. The comment that replaced them says why: “the app Josh depends on for the client, the price, the diary and the booking restarts many times a day, for a second or two each time, entirely by our own hand… land in one of those windows and the call simply fails.”
They now retry on a closed door — twice, with short waits, because a caller is on the line. Worst case is about 1.9 seconds of extra silence, not a failed booking. A caller would hear a pause. They would not lose the job.
And it writes down when it happens, in these words: recovered on attempt 2 — the first try hit a closed door (almost certainly a deploy restarting the app mid-call). So whether those two calls actually felt it is not a matter of opinion. It is a second grep, and it is in the run sheet below.
It still cannot be the main cause. Two calls against eleven cut-outs, and the Saturday session — 22 of the 37 calls — had not one commit to main in a twenty-four hour window either side. If the Saturday notes contain cut-outs, and they almost certainly do, something else is producing them.
This is the one that would really hurt: the phone bot is deployed by hand, and restarting it kills every call in progress. Josh’s brain changed nineteen times across these few days, so it is a fair thing to suspect.
None of them landed in a session. Zero changes during any of the three. The closest was Tuesday, and the last one that day was about eleven hours before the tester picked up the phone. Saturday had none for days in either direction.
The call log agrees. A phone bot killed mid-call never gets to file that call, so a killed call goes missing from the log rather than showing up as a short one. Saturday and Tuesday together are 37 logged calls against 36 written notes. Nothing is missing.
The 19 August report has the five that had happened by then, with dates: twice on 30 July, twice on 7 August, once on 11 August. None of them is near a session.
And they had never touched the phone anyway, which that report is blunt about: “the voice program uses about 790 MB, which makes it the second largest thing on the machine. The system kills the largest. It survived five times only because something else happened to be bigger each time.”
Four more have happened since, and one of them — 19 August at 8:51:45 PM — did take the phone bot down along with everything else. That was hours before the Wednesday session, not during it.
The tester was shooting at a moving target. Josh’s brain changed six times between the Saturday session and the Tuesday one, and thirteen more times between Tuesday and Wednesday.
So the thirty-six notes do not describe one Josh. They describe at least two, probably three — and nothing in the file says which note belongs to which. A fault written down on the Saturday may have been fixed by the Tuesday. A fault introduced on the Tuesday cannot appear in the Saturday notes at all.
That is the real cost of the restarts here, and it is why “which nights are these?” is not a tidiness question. Without it we will fix things that are already fixed and miss things that arrived last.
The hosting upgrade is still right — the machine has died nine times since 30 July and the phone survived on luck. What has changed is that it cannot be sold as the fix for the cut-outs, because on the biggest of the three sessions nothing restarted at all.
Both are in the phone bot. Both have been written about in its own code, before these tests happened.
Read the notes again and they have a shape. Josh does not go quiet at random moments. He stops while he is talking, and then apologises for it:
“Go ahead.” “I got it.” “I didn’t mean to interrupt you.” Those are not the words of a program whose audio dropped. They are the words of a program that believes the caller started speaking and is politely giving way. And it happens most on his longest sentences — the recording disclosure, a readback, a package pitch — which is exactly where there is most opportunity to be interrupted.
| Suspect 1 — a false barge-in | Suspect 2 — the number filter | |
|---|---|---|
| What happens | Something on the line — echo, breathing, road noise — is read as the caller talking, so Josh stops and waits | The filter that turns “$572” into words sits between his brain and his voice, and cuts the speech short |
| Why it is a suspect | How long he waits was lowered from 0.6s to 0.4s on Ken’s note about delay. The code comment predicted the bill: “lower it too far and Josh talks over a caller who paused mid-sentence.” Dil hit exactly that live on 16 August | The bot’s own code says so: “Josh has glitched mid-sentence on calls where the filter was running, including on a sentence with no number in it at all.” |
| How it is settled | The watchdog says barge_in=True | The watchdog says UNFINISHED … the caller never interrupted |
Every call now ends with a verdict written to the server’s log. It was added for precisely this — a greeting that died halfway on a live call and left no trace at all. It prints one of three things:
MUTE — not one frame of audio left the voice engine.N UNFINISHED … the caller never interrupted. This is the shape of
speech cut off at our end.OK … barge_in=true/falseNine calls already have a verdict waiting — the Wednesday 19 August session. Nobody has looked. That is the first thing to do and it takes a minute.
All three are read-only except the last, which is a switch built to be flipped.
pm2 logs austin-voice --lines 5000 --nostream | grep "silence-watchdog\] VERDICT" — every call since 19 August, one line each. If they say barge_in=true it is suspect 1; if they say UNFINISHED with no interruption it is suspect 2.pm2 logs austin-voice --lines 5000 --nostream | grep "stt-watchdog\] VERDICT" — the recogniser prints its own verdict on every call. A bad ear and a cut-off mouth look identical to a caller and are different faults.pm2 logs austin-voice --lines 5000 --nostream | grep "hit a closed door" — written whenever a lookup had to retry because the office app was restarting mid-call. Two pushes landed inside a live call on the 18th and the 19th; this says whether the caller heard them.AUSTIN_NUMBER_FILTER=off in /opt/austin-voice/.env, pm2 restart austin-voice, and have the tester repeat three scenarios. The switch was built to test the filter as a suspect and has never been used. Take it out again afterwards — without it prices are read as digits.How long Josh waits before deciding you have stopped has already moved twice on someone’s impression of one call, and both moves had a cost the other way. It is adjustable per phone number for exactly this reason — put the test line on a different value and run the same scenarios on both, rather than moving the number everybody is on and finding out afterwards.
Found while doing this, and it is why this investigation is thinner than it should be.
We keep the transcript of a call only if it booked. A transcript is filed on the CRM contact the booking created — no booking, no contact, nothing to attach it to, so it is dropped. Of the tester’s 47 calls, 24 did not book.
That is the wrong way round for testing. The calls worth reading are the ones that went wrong, and those are precisely the ones we throw away. The whole of GC’s thirty-six notes had to be analysed against the notes themselves, because our copy of what was said does not exist.
The recordings do survive. Telnyx keeps every call on GC’s number for a year, and the Calls tab in the portal plays them. So the audio is recoverable even where the transcript is not — and for a cut-out, the audio is better evidence anyway.
Keep the transcript of a failed call too, filed against the call rather than the contact. It is the same data we already send to the drift monitor and then discard. Small change, and it is the difference between reading the next thirty-six and guessing at them.