Handover · 20 August 2026
Written 2026-08-20 by the Claude on ClaudeAustin2026/Outcrop-Q. Companion to 2026-08-20-HANDOFF-dbclaude-memory-crash.md, which went to the DB Claude side and has already been built and shipped there. This is the same defect in Dean, and Dean is the one that actually caused the outage.
I could not reach your code. This session's GitHub access lists exactly three repos — Outcrop-Q, Claude-Austin, outcrop-orchestrator — and Dean is in none of them. So this is a document, not a pull request. Written to be picked up cold: you do not need to have seen any of the conversation behind it. Everything below was measured on the live server on 2026-08-19/20, and the fix described in part 3 is running in production in Outcrop-Q today, so you are copying something that works rather than designing from a description.
Dil passed this on the same evening it was written. Two things came back with it, and both are wider than Dean:
test/q-engine-guardrails.test.mjs runs where there is no systemd, so it can only exercise the watcher. The cgroup is the half that cannot be outrun between two polls, and the only place to prove it is the box that serves Q: cd /var/www/q-dashboard && node scripts/check-q-engine-guardrails.mjs
Look for systemd scope available: true. If it says false the ceiling still holds via the watcher — worth knowing, not worth panicking. Safe to run on the live server: it starts one short-lived allocator and kills everything it starts before it returns.
Also raised alongside it: access control on the dashboards themselves. That is a different problem from this one and should not be folded into it — see the two things NOT to do at the bottom.
The Claude that built DB Claude's guard read this document and found one of its own numbers wrong. Its reply, quoted because the reasoning is the useful part:
"I set
DBCLAUDE_HARD_MEM_MB=1400… my reasoning was 'every one of those was in the range that killed the box' — but that was the pre-swap box. With 4.6 GB of headroom now, the ceiling's job is to stop a runaway, not to stay under the old cliff. At 1400 I'd abort six of nine historically-completed sessions."
Nothing to deploy — three lines in /var/www/dbclaude-dashboard/.env and a restart, which is exactly why every limit here is an environment variable:
DBCLAUDE_HARD_MEM_MB=1600
DBCLAUDE_MIN_FREE_MB=400
DBCLAUDE_MAX_HISTORY_CHARS=240000
It is in this document's table and it was not built. It is the one that matters most for the symptom that started all of this, so it should not be quietly dropped:
A memory ceiling catches a turn that is growing. It does not catch a turn that is alive, small, and going nowhere — "Claude is a minute and a half, and it isn't doing anything." That one still spins forever.
Same pattern as the memory watchdog: a deadline, and an end with a sentence instead of a spinner. Recommended, and it is a contained change.
For the record, that Claude had already avoided Traps 2 and 3 independently while building, and sidestepped Trap 4 a different way — reading VmRSS from /proc/<pid>/status and walking task/*/children rather than parsing stat. Both routes are fine; the trap is only in parsing stat by splitting on whitespace.
A chat in a browser ran a 1.45 GB claude process on the box that answers Chad's and Ted's phones. The box had 1.35 GB free. The kernel killed it, PM2 went down with it, and all six apps and both public websites died at once.
The margin was 93 MB. The session that did it was Dean's.
The investigation was opened as a DB Claude problem, because that is where Beth was typing when it hung. It is not where the fatal process was.
| Where Beth was typing | the DB Claude dashboard |
| Which session the kernel killed | /root/.claude/projects/-root-dean/942aec5a-e6b8-4c05-a053-c584e34ec1a2 |
| Opened | 20:43 |
| Stopped being written | 20:51, at 154 KB |
| Killed | 20:51:45, anon-rss:1447296kB |
DB Claude was a victim. Dean was the cause. DB Claude has since been given a memory ceiling, a one-message-at-a-time rule, a size meter, a trimmed history and cleanup on tab close. Dean still has none of that, which means the app that actually took the phone line down is the one app still able to do it again.
That is the whole reason this document exists.
Server clock (CDT). Every line came off dmesg, journalctl, ps, pm2 or the filesystem.
| Time | What |
|---|---|
| 20:43 | The Dean session opens |
| 20:44:42 | systemd-resolved: Under memory pressure, flushing caches. The box is already starving, under two minutes after that session opened |
| 20:44–20:51 | Seven solid minutes of Under memory pressure, every ~2 seconds |
| 20:51 | The Dean session's .jsonl stops being written |
| 20:51:45 | Out of memory: Killed process 1037615 (claude) anon-rss:1447296kB |
| 20:51 | PM2's God Daemon restarts — taking q, dean, dbclaude, arlo, max and austin-voice with it |
| 20:52:14 | bookedsolidinspector.com answers 200 again, after 504s and 502s |
The seven minutes are the important part. That is not an outage, it is the machine visibly suffocating, and nobody could see it. Two people lived through the same seven minutes as two unrelated problems: Beth saw a chat that would not answer, and a separate investigation saw a website timing out after 71 seconds.
It was the ninth kill since 30 July — 4 on 30 July, 2 on 7 August, 1 on 11 August, 2 on 19 August. Every single victim was a claude process between 1.26 GB and 1.71 GB.
From the meeting recording, while it was happening:
claude process was the casualty.Clearing the chat fixed it immediately.
Beth had used these dashboards for weeks and had never once cleared a chat, because nothing ever told her to. No context indicator, no size warning, no cap, no rotation. Ken and Dil both diagnosed it correctly in the room in three minutes; the software never hinted at it in weeks. Nobody was careless.
/swapfile2, /etc/fstab updated). Headroom went from 1,354 MB → 4,606 MB, against a largest-ever session of 1.71 GB./usr/local/bin/memory-watch.sh, cron every 5 minutes, logging to /var/log/memory-watch.log. Source in Outcrop-Q at scripts/memory-watch.sh. It only watches — nothing that kills or restarts, because something that restarts things could take the phone line down at 2 AM.sysstat was already installed and nobody knew: sar -r -f /var/log/sysstat/sa<DD> has days of history.The crashes should now stop. The slowness will not — swap is disk, so a very large chat still thrashes. The phone bot has oom_score_adj -1000 and is exempt from the killer, but nothing exempts it from swap thrashing, and for a voice agent latency is the product.
A chat in a browser must never be able to kill PM2 and take a phone line with it. Put a memory ceiling on the claude process Dean spawns, so a runaway turn fails itself and shows the user a sentence, instead of the kernel picking a victim elsewhere on the box.
Everything else on this list is prevention. This is containment, and it is the difference between "Dil sees an error" and "Chad's calls stop".
Beth sent the same message twice and then "hello" twice into a chat that was already struggling. Ken's read was mechanically right: each send starts another process carrying the same huge context, so the second send makes it strictly worse.
Disable the send button while a turn is running — and then enforce it server-side anyway. A disabled button is bypassed by refreshing the page.
A size meter, a token readout, or simply "this chat is long — start a new one" past a threshold. Beth had no way to know. Neither does anyone else.
If Dean flattens the whole conversation into every turn — Q did, DB Claude did — then each turn of a long chat starts bigger than the last. Trim the OLDEST turns, never the newest message, and tell the model what was dropped so it asks instead of guessing.
Before this fix, closing the browser left the process running to finish an answer nobody would ever read.
/root/.claude/projects/-root-dean/ is worth an ls -lat. On the DB Claude side there were 105 session directories, ~150 MB, and six new sessions appeared in four minutes during the recovery.
api/q-pro/guardrails.js in ClaudeAustin2026/Outcrop-Q is a working version of all of the above, in production, with 24 automated checks. If you can get read access to that repo, copy it. If not, the traps below are the part that costs days.
| Setting | Value | Why |
|---|---|---|
| Soft warn | 900 MB | logs and warns; the turn continues |
| Hard stop | 1600 MB | ⚠️ not 1400. Six of the nine historical kills were above 1400 MB, so a 1400 ceiling would kill turns that used to succeed. With 4 GB of swap the box absorbs 1.7 GB, so the ceiling's job is "stop a runaway", not "stay under the old cliff" |
| Refuse to start below | 400 MB available | better a clear "not now" than joining the queue for the OOM killer |
| Transcript budget | 240,000 chars | ≈ 60k tokens of history |
| Poll | every 2 s | |
| Wall-clock timeout | 5 min | a turn that is alive but going nowhere must end with a sentence, not a spinner |
Every one of them is an environment variable. Tuning must never need a code change and a deploy — the deploy restarts PM2, which is its own small outage.
MemoryHigh does not warn. It throttles.This was a live bug, caught by a check on the real server, and it is the most expensive thing in this document.
Setting MemoryHigh at the soft limit put the cgroup under heavy reclaim at 120 MB, so a process allocating 20 MB every 50 ms never reached the ceiling at all — the watcher had nothing to fire on, and the turn sat there alive and crawling.
That is strictly worse than having no protection. It converts "stopped with an explanation" into "hangs forever" — the exact symptom the whole investigation started from: a chat that spun for a minute and a half.
The order that works:
your own watcher stops the turn at the ceiling, with a message a human can read
MemoryHigh ceiling + 100M — back-pressure ONLY above that, if the watcher was slow
MemoryMax ceiling + 200M — the kernel's absolute wall
Each layer backstops the one before it. None pre-empts it.
NODE_OPTIONS is stripped by the SDKThe obvious --max-old-space-size route is a silent no-op. It looks configured and does nothing. Which is exactly the kind of fix that gets shipped, believed, and then fails.
claude on this server is a native binary, not a node scriptSo node flags passed through executableArgs arrive as CLI arguments to claude and are rejected, or worse, ignored. The ceiling has to be external — a cgroup and a watcher — not a flag.
On the systemd path your direct child is systemd-run, which is a few hundred KB. Measuring only the direct child reports a comfortable 0 MB while 1.5 GB sits one level down.
Two details in doing that: read /proc directly rather than shelling out to ps (this runs every couple of seconds for the life of every turn), and parse /proc/<pid>/stat from the last ) — field 2 is the executable name, it is wrapped in parentheses, and it can contain spaces and parentheses of its own. Splitting the line on whitespace is the classic bug and it mis-reads the parent pid.
MemAvailable, not freeos.freemem() excludes reclaimable page cache, so a healthy Linux box looks alarming and a starving one looks fine. /proc/meminfo → MemAvailable is the number the kernel itself uses.
systemd-run on PATH does not mean systemd is running/run/systemd/system existing is the canonical test. The binary is present and useless inside containers.
Q's suite deliberately starts a process designed to eat memory and proves the ceiling stops it. Then the code was broken four different ways on purpose to confirm the tests caught every one. A test that never fails is not a test — and on this particular feature, a test that passes against a silent no-op is exactly how Trap 2 ships.
Do not delete the saved chat history, and do not remove the chat sidebar. It was proposed as a fix on the DB Claude side and it is the wrong target.
If the real concern is who can read those chats, that is an access-control question — a login — not a deletion question. Deleting old chats does not stop anyone reading tomorrow's.
dmesg -T | grep -i "killed process" # the nine kills
journalctl --since "2026-08-19 20:40" --until "2026-08-19 21:05" --no-pager
ls -lat /root/.claude/projects/-root-dean/ | head # session 942aec5a, 20:43 → 20:51
free -h && swapon --show # the swap that was added
cat /proc/$(pgrep -f "austin-voice.*bot.py")/oom_score_adj # must be -1000
tail -40 /var/log/memory-watch.log # the monitor
A browser chat window quietly ran a 1.5 GB process on the machine that answers a real client's phone, and there was no way for anyone to see it happening. The system had no gauge, no cap and no warning — and it was 93 MB away from being fine.
Dean is the last of the three that can still do it.