Split out of #11, where this sat since 2026-09-24 as "a PTY exit goes unobserved (Linux)". Rewritten: the first version of this issue blamed the call site and was wrong. What follows is what survived measurement.
The symptom
On the CI for #10:
(fail) v2 pty HttpApi > serves location-wrapped PTY routes and retains exited sessions [21256.87ms]
- "status": "exited", "exitCode": 4
+ "status": "running", "pid": 6907, …
A sh -c "exit 4" polled every 50ms for 20 seconds and never left running. It did not reproduce on a re-run and passes locally. Since only the onExit event ever sets that status, the exit was not reported at all.
What it is not
Not a race between the spawn and the listener attach in Pty.create. That was the original reading here and it does not survive measurement: over 4901 trials, including pads walking the fiber's op budget across the boundary (maxOpsBeforeYield is 2048), there were 0 microtask and 0 macrotask crossings between yield* Effect.sync(...) and the statements after it. Effect.sync invokes its continuation inline, so the spawn and both push(proc.onData(…), proc.onExit(…)) calls already run in one JS turn, and nothing can fire in between. Two independent reviews reached the same conclusion, one of them also unable to reproduce any loss on dev over repeated runs.
What it is, as far as the evidence goes
bun-pty can lose the exit before any listener could exist. _startReadLoop() is the last statement of the Terminal constructor (bun-pty/src/terminal.ts:175), it is async but has no await before its first bun_pty_read, and its emitter keeps nothing for latecomers (interfaces.ts: fire walks the listeners registered at that instant). So the first read, and any event it triggers, happens inside the constructor, before spawn returns:
n === -2 (CHILD_EXITED) → _onExit.fire(...) to zero listeners, then break. The loop is over, so nothing reports it again and the session stays running for its whole life.
n > 0 → the first output chunk is fired to zero listeners and dropped.
n < 0 (read error) → the loop breaks without firing onExit at all. Even a correctly attached listener gets nothing. Same stranded session.
kill() fires _onExit synchronously with a fabricated exitCode: 0.
Honest about reachability: I could not make the constructor window fire in 1500 clean sequential /bin/true spawns with listeners attached synchronously (0 lost). A first probe that did show ~350 losses in 1500 was invalid — it never released the PTY handles, so it was exhausting them and hitting the n < 0 branch instead. A review reported 6 of 4000 first reads returning -2 at the FFI level and 1 lost exit in 2500 end to end; neither reproduced here, and the sequential shape of that probe is the one that leaks handles, so I am not treating those numbers as established.
The n < 0 branch is the more promising lead for the CI symptom, because that is exactly what resource pressure produces, and a loaded runner with many PTYs open across a suite is where this appeared. That is consistent with the invalid probe above: under handle exhaustion, exits stopped being delivered entirely.
What would fix it
Nothing at the call site, and nothing in packages/core/src/pty/pty.bun.ts either: Terminal exposes neither the read-loop state nor bun_pty_get_exit_code, so the shim cannot ask afterwards whether an exit was missed. The options are
- a vendored or upstream bun-pty change — defer
_startReadLoop() past a microtask, or give its EventEmitter replay until the first subscriber, and make the n < 0 branch report an exit instead of going silent (packages/core/script/fix-node-pty.ts is a precedent for a post-install fixup), or
- reconciliation in
Pty: notice that a process recorded as running is gone and settle its status without an event.
Both are design decisions rather than small fixes, which is why this stays open. #30 does not close it — it fixes the separate ordering defect in #32 that the investigation turned up.
Split out of #11, where this sat since 2026-09-24 as "a PTY exit goes unobserved (Linux)". Rewritten: the first version of this issue blamed the call site and was wrong. What follows is what survived measurement.
The symptom
On the CI for #10:
A
sh -c "exit 4"polled every 50ms for 20 seconds and never leftrunning. It did not reproduce on a re-run and passes locally. Since only theonExitevent ever sets that status, the exit was not reported at all.What it is not
Not a race between the spawn and the listener attach in
Pty.create. That was the original reading here and it does not survive measurement: over 4901 trials, including pads walking the fiber's op budget across the boundary (maxOpsBeforeYieldis 2048), there were 0 microtask and 0 macrotask crossings betweenyield* Effect.sync(...)and the statements after it.Effect.syncinvokes its continuation inline, so the spawn and bothpush(proc.onData(…), proc.onExit(…))calls already run in one JS turn, and nothing can fire in between. Two independent reviews reached the same conclusion, one of them also unable to reproduce any loss ondevover repeated runs.What it is, as far as the evidence goes
bun-pty can lose the exit before any listener could exist.
_startReadLoop()is the last statement of theTerminalconstructor (bun-pty/src/terminal.ts:175), it isasyncbut has noawaitbefore its firstbun_pty_read, and its emitter keeps nothing for latecomers (interfaces.ts:firewalks the listeners registered at that instant). So the first read, and any event it triggers, happens inside the constructor, beforespawnreturns:n === -2(CHILD_EXITED) →_onExit.fire(...)to zero listeners, thenbreak. The loop is over, so nothing reports it again and the session staysrunningfor its whole life.n > 0→ the first output chunk is fired to zero listeners and dropped.n < 0(read error) → the loopbreaks without firingonExitat all. Even a correctly attached listener gets nothing. Same stranded session.kill()fires_onExitsynchronously with a fabricatedexitCode: 0.Honest about reachability: I could not make the constructor window fire in 1500 clean sequential
/bin/truespawns with listeners attached synchronously (0 lost). A first probe that did show ~350 losses in 1500 was invalid — it never released the PTY handles, so it was exhausting them and hitting then < 0branch instead. A review reported 6 of 4000 first reads returning-2at the FFI level and 1 lost exit in 2500 end to end; neither reproduced here, and the sequential shape of that probe is the one that leaks handles, so I am not treating those numbers as established.The
n < 0branch is the more promising lead for the CI symptom, because that is exactly what resource pressure produces, and a loaded runner with many PTYs open across a suite is where this appeared. That is consistent with the invalid probe above: under handle exhaustion, exits stopped being delivered entirely.What would fix it
Nothing at the call site, and nothing in
packages/core/src/pty/pty.bun.tseither:Terminalexposes neither the read-loop state norbun_pty_get_exit_code, so the shim cannot ask afterwards whether an exit was missed. The options are_startReadLoop()past a microtask, or give itsEventEmitterreplay until the first subscriber, and make then < 0branch report an exit instead of going silent (packages/core/script/fix-node-pty.tsis a precedent for a post-install fixup), orPty: notice that a process recorded as running is gone and settle its status without an event.Both are design decisions rather than small fixes, which is why this stays open. #30 does not close it — it fixes the separate ordering defect in #32 that the investigation turned up.