Investigation · started Thursday 20 August 2026

The call records behind GC’s thirty-six notes

Dil asked for the call records first, and he was right to. Six hundred and twenty-two calls are on file. Isolating the tester’s gives the sessions, the times and the durations — and it kills the leading suspect for the cut-outs. The releases are not doing it. What is left is two candidates, one command that separates them, and a switch built to settle exactly this that has never been thrown.

47calls from GC’s tester
3testing sessions
2calls with the app restarted under them
9calls with a watchdog verdict waiting
The short version

The tester is one number, and it has rung Josh 47 times. Three sessions: Saturday 15 August, Tuesday 18 August and Wednesday 19 August. Twenty-three of the 47 booked.

The release theory cannot carry it — but it is not zero, and my first pass said it was. Correcting that here: a first count compared timestamps as text, and six of the commits carry a +08:00 offset. Redone on real instants, two pushes landed inside a live call. Both are below, with the calls they landed in.

It still cannot be the main cause. Two calls, against eleven cut-outs in the notes — and the Saturday session, 22 of the 37 calls, had not one commit to main in a twenty-four hour window either side.

Two suspects are left, and they are both in the phone bot. A false barge-in — Josh hearing something and yielding mid-sentence — or the number filter mangling his speech on the way out. One command tells them apart, and the answer for nine of those calls is sitting in the server’s log right now.

The sessions

One caller, three evenings. Times are Central, the way GC works.


SessionCallsBookedWindowMedian lengthLongest
Fri 14 Aug1017:588m 07s8m 07s
Sat 15 Aug221013:45 – 19:085m 04s9m 13s
Tue 18 Aug15921:18 – 22:554m 28s6m 12s
Wed 19 Aug9406:25 – 23:093m 43s8m 27s

The thirty-six notes are almost certainly the 15th and the 18th. Those two sessions are 37 calls, and one of them lasted twelve seconds — which is a redial, not a scenario. This needs one line back from the tester rather than a guess, because the whole analysis below hangs on which nights are being described. If the notes are the 18th and 19th instead, one of the two suspects changes.

Two numbers worth having beside the notes

The tester books 49% of the time. Real callers on the same line book 58%. That gap is expected — several scenarios are written to hang up once a price is given — but it is the baseline any future run should be read against.

Four calls ended in under thirty seconds — 10, 10, 12 and 14 seconds. Call 22 in the notes is the one that reads “I never finished the call because I was so frustrated”, and a call that short is somebody hanging up, not a conversation that failed.

Did our own restarts hit these calls?

Dil’s question, and it has three separate answers because there are three separate kinds of restart.


“A restart” covers three different things here, and only one of them touches a call. Separating them is most of the answer.

Kind of restartWhat it restartsDid it hit these calls?
A release — every push to mainThe office app only. pm2 restart q. It does not touch the phone botYes, twice — both inside a live call
A phone deploy — copied by handThe phone bot itself. Kills any call in progressNo. Not one during any session
The machine running out of memoryWhatever it picks. Can take everythingNo. All of them missed the sessions

1. The releases — two landed inside a live call

The first pass on this said zero, and it was wrong. It compared timestamps as text, and six of the commits are written with a +08:00 offset, so they sorted into the wrong place. Redone properly:

SessionCallsPushes during itLanded in a live call?
Sat 15 Aug220 — and none for 24h either side—
Tue 18 Aug151, at 03:16:44ZYes — 24 seconds before a call ended
Wed 19 Aug91, at 03:33:33ZYes — about two thirds through a call

What a release actually does to a call in progress. It does not restart the phone bot, so Josh keeps talking. It restarts the app he fetches the price list, the diary, the caller lookup and the booking from. For a second or two those are a closed door.

⭐ This exact question was answered in the code on 7 August, and it left a log line behind

Every one of those lookups was a single attempt until then. The comment that replaced them says why: “the app Josh depends on for the client, the price, the diary and the booking restarts many times a day, for a second or two each time, entirely by our own hand… land in one of those windows and the call simply fails.”

They now retry on a closed door — twice, with short waits, because a caller is on the line. Worst case is about 1.9 seconds of extra silence, not a failed booking. A caller would hear a pause. They would not lose the job.

And it writes down when it happens, in these words: recovered on attempt 2 — the first try hit a closed door (almost certainly a deploy restarting the app mid-call). So whether those two calls actually felt it is not a matter of opinion. It is a second grep, and it is in the run sheet below.

It still cannot be the main cause. Two calls against eleven cut-outs, and the Saturday session — 22 of the 37 calls — had not one commit to main in a twenty-four hour window either side. If the Saturday notes contain cut-outs, and they almost certainly do, something else is producing them.

2. The phone bot itself — never restarted during a session

This is the one that would really hurt: the phone bot is deployed by hand, and restarting it kills every call in progress. Josh’s brain changed nineteen times across these few days, so it is a fair thing to suspect.

None of them landed in a session. Zero changes during any of the three. The closest was Tuesday, and the last one that day was about eleven hours before the tester picked up the phone. Saturday had none for days in either direction.

The call log agrees. A phone bot killed mid-call never gets to file that call, so a killed call goes missing from the log rather than showing up as a short one. Saturday and Tuesday together are 37 logged calls against 36 written notes. Nothing is missing.

3. The machine running out of memory — and the phone kept getting lucky

The 19 August report has the five that had happened by then, with dates: twice on 30 July, twice on 7 August, once on 11 August. None of them is near a session.

And they had never touched the phone anyway, which that report is blunt about: “the voice program uses about 790 MB, which makes it the second largest thing on the machine. The system kills the largest. It survived five times only because something else happened to be bigger each time.”

Four more have happened since, and one of them — 19 August at 8:51:45 PM — did take the phone bot down along with everything else. That was hours before the Wednesday session, not during it.

⚠️ The restarts did affect the test, just not the way the question assumes

The tester was shooting at a moving target. Josh’s brain changed six times between the Saturday session and the Tuesday one, and thirteen more times between Tuesday and Wednesday.

So the thirty-six notes do not describe one Josh. They describe at least two, probably three — and nothing in the file says which note belongs to which. A fault written down on the Saturday may have been fixed by the Tuesday. A fault introduced on the Tuesday cannot appear in the Saturday notes at all.

That is the real cost of the restarts here, and it is why “which nights are these?” is not a tidiness question. Without it we will fix things that are already fixed and miss things that arrived last.

What this does not clear

The hosting upgrade is still right — the machine has died nine times since 30 July and the phone survived on luck. What has changed is that it cannot be sold as the fix for the cut-outs, because on the biggest of the three sessions nothing restarted at all.

What is left, and how to tell them apart

Both are in the phone bot. Both have been written about in its own code, before these tests happened.


Read the notes again and they have a shape. Josh does not go quiet at random moments. He stops while he is talking, and then apologises for it:

“Go ahead.” “I got it.” “I didn’t mean to interrupt you.” Those are not the words of a program whose audio dropped. They are the words of a program that believes the caller started speaking and is politely giving way. And it happens most on his longest sentences — the recording disclosure, a readback, a package pitch — which is exactly where there is most opportunity to be interrupted.

Suspect 1 — a false barge-inSuspect 2 — the number filter
What happensSomething on the line — echo, breathing, road noise — is read as the caller talking, so Josh stops and waitsThe filter that turns “$572” into words sits between his brain and his voice, and cuts the speech short
Why it is a suspectHow long he waits was lowered from 0.6s to 0.4s on Ken’s note about delay. The code comment predicted the bill: “lower it too far and Josh talks over a caller who paused mid-sentence.” Dil hit exactly that live on 16 AugustThe bot’s own code says so: “Josh has glitched mid-sentence on calls where the filter was running, including on a sentence with no number in it at all.”
How it is settledThe watchdog says barge_in=TrueThe watchdog says UNFINISHED … the caller never interrupted
⭐ There is already a watchdog on this, and it has been running since 19 August

Every call now ends with a verdict written to the server’s log. It was added for precisely this — a greeting that died halfway on a live call and left no trace at all. It prints one of three things:

Nine calls already have a verdict waiting — the Wednesday 19 August session. Nobody has looked. That is the first thing to do and it takes a minute.

What to run, in this order

All three are read-only except the last, which is a switch built to be flipped.


1 — read the verdicts we already havepm2 logs austin-voice --lines 5000 --nostream | grep "silence-watchdog\] VERDICT" — every call since 19 August, one line each. If they say barge_in=true it is suspect 1; if they say UNFINISHED with no interruption it is suspect 2.
2 — check the ear at the same timepm2 logs austin-voice --lines 5000 --nostream | grep "stt-watchdog\] VERDICT" — the recogniser prints its own verdict on every call. A bad ear and a cut-off mouth look identical to a caller and are different faults.
2b — and whether our own releases were felt on the linepm2 logs austin-voice --lines 5000 --nostream | grep "hit a closed door" — written whenever a lookup had to retry because the office app was restarting mid-call. Two pushes landed inside a live call on the 18th and the 19th; this says whether the caller heard them.
3 — only if it says UNFINISHED: run the A/B that already existsPut AUSTIN_NUMBER_FILTER=off in /opt/austin-voice/.env, pm2 restart austin-voice, and have the tester repeat three scenarios. The switch was built to test the filter as a suspect and has never been used. Take it out again afterwards — without it prices are read as digits.
And do not tune the barge-in by feel if it turns out to be suspect 1

How long Josh waits before deciding you have stopped has already moved twice on someone’s impression of one call, and both moves had a cost the other way. It is adjustable per phone number for exactly this reason — put the test line on a different value and run the same scenarios on both, rather than moving the number everybody is on and finding out afterwards.

A gap in our own records, worth closing

Found while doing this, and it is why this investigation is thinner than it should be.


We keep the transcript of a call only if it booked. A transcript is filed on the CRM contact the booking created — no booking, no contact, nothing to attach it to, so it is dropped. Of the tester’s 47 calls, 24 did not book.

That is the wrong way round for testing. The calls worth reading are the ones that went wrong, and those are precisely the ones we throw away. The whole of GC’s thirty-six notes had to be analysed against the notes themselves, because our copy of what was said does not exist.

The recordings do survive. Telnyx keeps every call on GC’s number for a year, and the Calls tab in the portal plays them. So the audio is recoverable even where the transcript is not — and for a cut-out, the audio is better evidence anyway.

The recommendation

Keep the transcript of a failed call too, filed against the call rather than the contact. It is the same data we already send to the drift monitor and then discard. Small change, and it is the difference between reading the next thirty-six and guessing at them.