Daily build report · Saturday 8 August 2026

The fallback works. It is not switched on for a single client.

A question open since 30 July got answered today, on two real phone calls: when our end is down, a client's callers can be handed to that client's own phone in four thousandths of a second. Everything needed to give Chad and Ted that protection is now built and tested. None of it is switched on — and all three things standing in the way are human, not technical.

4mshandover, measured
2proving calls
3tools built
17new checks
3blockers, all human
The short version, for anyone not reading the rest

Once a client moves their phone number to us, they can't undo it in a hurry — un-porting takes days. So if we go dark, they are stuck with us being dark. We owed them an answer that works with nobody at our end awake.

We now have one, and it is proven rather than assumed: the instruction telling the phone company what to do with their number is stored at the phone company, not on our computers, so it still runs when ours are off. It says "send the audio to Josh; if that fails, ring the inspector." Today we made it fail on purpose and watched it ring.

What is not done: no client number has that instruction on it yet, we don't have Chad's or Ted's mobile number, and nobody is assigned to the call-backs Josh promises when he gets stuck. Those are the next three things, and two of them are conversations, not code.

What we set out to prove

Yesterday's attempt at this hit a wall. Today's started by going around it.


1

The obvious protection is not available to us, and that turned out to be the clue

The natural way to protect a ported number is call forwarding at the phone company: if our end doesn't answer, send the call somewhere else. On 7 August Telnyx refused it outright — error 10015, no automatic forwarding on any number attached to a voice application. Every one of our client numbers is on one, so that door is shut for all of them.

But the error message pointed somewhere. It said to use the call control functionality instead. The instruction sheet Telnyx follows for each number is already stored on their computers — which is the one place that keeps working when ours don't. If that sheet could carry on to a second instruction after the first one failed, the whole problem was solvable.

Nobody had tested whether it does. Everything written about this said it "should". Eleven silent faults reached production in a single day last month on the strength of "should", so it got tested instead.

10015 — architectural, not a bad setting

The two calls that proved it

Both made by Dil, on the spare test line, with nothing of Chad's or Ted's touched.


2

Call one — a stream pointed at a host that cannot exist

A test instruction sheet was built that sends the audio to this-host-does-not-exist.invalid, then says a sentence, then dials a second number. The middle sentence is the whole point: without it, a silent call could mean either "the sheet stopped" or "the sheet carried on and the dial failed", and those need completely different fixes.

Dil rang it at 03:19. Telnyx's own event log: the stream failed at 41.157, and the speech began at 41.161.

Four milliseconds. And from the caller's side — answered at 41.098, voice at 41.384 — under three tenths of a second of silence. That number decides whether the design is usable at all. Had a failed stream taken the eight or ten seconds it might have, no caller would have stayed on the line and the whole approach would have been dead regardless of what the log said.

sheet continues past a failed stream<0.3s dead air
3

Call two — the dial half, proven without waking anybody at 2am

Call one proved the sheet carries on. It did not prove the second leg actually connects a human. Testing that needs a phone somebody will answer, and it was the middle of the night in the Philippines — Dil's words: "I cannot ring Beth, she's asleep now."

The way round it: point the fallback at one of our own numbers, so the second leg rings a line Josh answers. Nobody gets woken, and the chain is still fully exercised — failed stream, spoken handover, outbound dial, second call connecting, a voice on the other end.

Dil's report afterwards was two words long: "Yes Josh talked." The Telnyx log showed the second leg being placed and answered. The full chain works end to end.

full chain provennobody woken

How it actually works

Two drawings. The first is why any of this is possible; the second is every path a call can take and what fires on it.


Drawing 1Where each piece lives
The caller rings the number THE PHONE COMPANY The client's ported number e.g. Chad's line in Pearland THE INSTRUCTION SHEET one per client, held by the phone company 1 · send the audio to Josh 2 · if that fails, ring the inspector Line 2 runs — they dial out only when line 1 could not connect EVERYTHING RIGHT OF THIS LINE CAN BE DOWN AT ONCE OUR SERVER Josh's ears and mouth holds the whole live call austin-voice The main application prices, availability, bookings writes to the CRM SUPPLIERS Deepgram Claude ElevenLabs The inspector's own mobile rings for 20 seconds live audio
The phone company holds both the number and the sheet that says what to do with it. Our server is one destination the sheet points at — a destination that can vanish without taking the sheet with it. That is the entire reason a fallback is possible. It is also why nothing at all works when the phone company is the thing that's down.
Drawing 2One call, every fork, and what fires
THE CALL WHAT FIRES, AND WHAT THE CALLER GETS A caller rings the number the client ported to us The sheet for that number is opened it is stored at the phone company — see drawing 1 Line 1: can they reach our server? the audio stream opens, or it doesn't yes Josh answers Whose company is this number? the number dialled decides whose prices he quotes recognised Hears the caller Works out what to say Says it every turn Saves the booking to the CRM Booked work order + call transcript filed on the contact no 4 ms later Josh is skipped. The sheet moves to line 2. measured 8 Aug: under 0.3 s of quiet for the caller Caller hears: “One moment please, I’m connecting you now.” The inspector’s own mobile rings — 20 s he answers it himself, or it goes to his voicemail unknown Josh takes a message and nothing else no price, no time promised, no booking stops a misheard digit becoming another firm’s prices silent No answer from the model for 12 s. Twice. the brain watchdog calls it dead and stops waiting Josh: “I’ll have someone call you straight back.” The call is flagged for a call-back ⚠ the inspector’s mobile does NOT ring — the phone company saw the call answered, so line 2 never runs. refused We were mid-restart — try again at 0.4 s, 1.5 s nothing was delivered, so a retry cannot double-book Timed out AFTER sending — never retried it may already be saved. A second attempt would be a second booking, and a duplicate is worse than a delay. If the PHONE COMPANY is the thing that’s down, none of this picture runs. The number, the instruction sheet and the fallback dial all live there. Nothing in our control covers it — the honest answer is a second phone company holding the same numbers, which is a business decision, not a fix we can ship.
the call completes degraded — Josh still handles it a person takes over left for a human, or not covered
The fork at the top and the fork in the middle look similar and are not. When our server is unreachable, the phone company never gets an answer and moves to line 2 — the mobile rings. When Josh answers and then the model goes quiet, the phone company has already seen a successful call, so line 2 will never run no matter how long the silence lasts. That is why the watchdog has to live inside Josh rather than in the instruction sheet.

What got built today

Three tools and a section of website. Two are ready to run; one is deliberately not published.


4

The generator — everything needed to protect a client, in one command

It reads the backup number the inspector gave us, builds his instruction sheet with his own mobile in it, stores it at the phone company, and points his number at it. One command per client instead of three fiddly steps in a web console.

It refuses more than it does. It changes nothing unless told twice — a dry run by default, and a second explicit flag before it will touch a real client's line. It rejects a backup number that is the same as the client's own number, which would make the call ring itself in a loop and be billed for every leg. It reads the sheet back from the phone company and compares it before pointing anyone's number at it, because a sheet that uploaded badly takes that number off the air completely. And a client with no backup number on file is reported in capital letters as NOT PROTECTED rather than quietly skipped — an unprotected client who looks protected is worse than one you know about.

One architectural decision worth Ken seeing. Our written rule was "one sheet shared by every number, nothing per-client to create." That was the right call and it was made before any of this was a requirement. A shared sheet has no way of knowing which client's call it is handling, so there is nowhere to put that inspector's own mobile. The moment each client falls back to his own phone, it has to be one sheet per client. The cost is stated rather than discovered: adding a client is now three steps instead of one, and a half-finished one is invisible until an outage — so there is a command that lists exactly who is and isn't protected, to be run after every signup.

built & dry-run clean17 checksnot run on anybody
5

Something that remembers to put things back

Two things were quietly left switched in the last fortnight, and neither raised an error anywhere. The office number stopped forwarding to Beth while our own notes still said it was forwarding — anyone ringing the number printed on the website reached nobody, and we found it by accident while looking at something else. And a test token sat live in the settings from early July to 7 August, five weeks, leaving a door open to anyone who guessed the address.

Neither was carelessness. Both were switched deliberately, for a good reason, by someone who fully intended to switch them back. They were temporary by intent and permanent by default, because nothing remembered.

There is now a written register — you add a line before you flip a switch, with a date — and a checker that fails the moment that date passes. It also looks at the live state: test tokens left armed, a client number pointing at nothing, the office number not forwarding, the test line left attached to a client's setup, and instruction sheets approaching expiry. When it can't reach the phone company it says so out loud instead of reporting all clear — that false comfort is exactly what the old note about the office number was.

register + checker livenpm run put-it-back
6

The signup form now asks for the backup number — and says what it costs

The whole design depends on knowing which phone to ring, and there was nowhere to tell us. There is now, on the live signup form, with a plain note underneath rather than a bare field: that this phone may ring at two in the morning, that it should be a phone the inspector actually answers, and that leaving it blank is a real option if he would rather callers reached his voicemail than his bedside.

That last part matters. This is his decision to make and not ours to assume — and asking it at signup turns it into a promise he agreed to, rather than a surprise he discovers during an outage.

live on the signup form
7

The outage section for the public website — written, and deliberately not published

Dil asked for the contingency plan to go on bookedsolidinspector.com. It is written: three plain panels after "What Booked Solid Inspector Can't Do" — what happens when our end is down, what happens when Josh gets stuck, and a flat admission that if the phone company goes down nothing we do helps.

It is sitting in the repository unpublished, and it should stay there until Ken or Beth says otherwise. The first panel tells a prospect their calls ring their own phone when we cannot answer. That is true of the design and true of no live number today. The second promises a human rings back, and nobody is assigned to that.

Publishing it would be the same shape of failure as the three clients who were never real and the price table that was never charged: something believed for months because it was written down somewhere authoritative. The section is good and it should go live — the week after it becomes true, not the week before.

needs Ken's or Beth's rulingin the repo, not on the site

What went wrong on the way

Recorded because the near-misses are the useful part.


8

I took the Hamming test line off the air with an ambiguous sentence

To prove the second leg without waking anybody, I told Dil to point the fallback at 888-347-2042. He reasonably read that as "attach that number to the test setup", did so, and the next call to it returned "an application error has occurred." That number is the line Hamming runs its automated tests against — while it was moved, those tests would all have failed.

It was back within minutes and the proving call went ahead. The lesson isn't "be careful"; it is that a switch made for a test needs writing down at the moment it is made. That is precisely what the register built today is for, and its first entries are the three things switched during this session — all three now back.

restored same session
9

The test rig would happily have dialled itself in a loop

Asked which number to use as the fallback, Dil reached for the test line — the number in front of him, and the natural thing to reach for. That makes the call dial itself: in, stream fails, dial the same number, in again, with every leg billed.

The rig now refuses it outright, and so does the generator, comparing on the last ten digits so a different way of writing the same number can't slip past. In a test it wasted a few minutes. Baked into a live client's sheet it would have been a billing incident.

refused in both places

What's still open

In order. The first two are conversations, not builds.


Three things, and only one of them is code

  1. Ask Chad and Ted which phone should ring. Nothing can be switched on without an answer, and it is genuinely theirs to give — this phone may ring at 2am. Worth asking now, in a calm moment, rather than during an outage. "Leave it blank" is an acceptable answer.
  2. Decide who handles the call-backs. When Josh gets stuck he promises a human will ring straight back, and the call gets flagged. Today nothing watches for that flag. The mechanical half is live and working; the promise is the part most likely to break, and it needs a person's name against it, not a script.
  3. Then run the generator, and make one test call per number. Roughly a minute of work per client once the answers exist. The test call is not optional — a mis-pointed number answers with an error and stays broken until somebody rings it.

And one thing that cannot be fixed from here

  1. If the phone company itself goes down, nothing we build helps. The number, the instruction sheet and the fallback dial all live there. The only real answer is a second phone company holding the same numbers, which is a cost and a business decision. Raising it so it's a known and accepted risk rather than a discovered one.
The one number to remember from today

Four milliseconds, and under three tenths of a second of silence for the caller. That is the gap between our server dying and the client's own phone starting to ring. It is the difference between an outage the caller notices and an outage they don't — and it is measured on a real call, not estimated.