Found by accident, and it is not a Windows-runner story: it reproduces on a four-core Linux box in this container, with an unmodified checkout.
The measurement
packages/opencode/test/session/prompt.test.ts:2011, run on its own (bun test test/session/prompt.test.ts -t "cancel interrupts loop queued behind shell" --timeout 30000), on dev at 37a3827c51 with no local changes:
run 1: ^ this test timed out after 30000ms. 0 pass 1 fail
run 2: 1 pass
run 3: 1 pass
run 4: 1 pass
One failure in four, and it is the test's own 30s ceiling rather than an assertion. Two neighbouring variants of an unrelated log line gave 0 failures in 4 and 2 in 3 — all small samples of the same flake, which is how it was noticed: I attributed a 30s hang to a line I had just moved, and the control run without that line hangs too. That attribution is retracted on #40.
What it is waiting for
const sh = yield* prompt.shell({ sessionID: chat.id, agent: "build", command: "sleep 30" }).pipe(Effect.forkChild)
yield* waitForBusy(chat.id)
const loop = yield* prompt.loop({ sessionID: chat.id }).pipe(Effect.forkChild)
yield* Effect.sleep(50)
yield* prompt.cancel(chat.id)
const exit = yield* Fiber.await(loop) // asserts "User aborted the command"
…
yield* Fiber.await(sh) // unbounded
The command is sleep 30 and the ceiling is 30s, so a cancel that does not actually end the shell is indistinguishable from the child simply running to completion. Since a failing assertion on the loop would report as an assertion, the timeout means one of the two Fiber.awaits never returned — and the test cannot say which, because the last one is unbounded. #28 gave the two sibling tests in this file named, bounded waits for exactly this reason; this one did not get the same treatment. Doing that is the cheap first step and would make the next occurrence say something.
The spawner is not where it is stuck
Worth recording so nobody repeats it. I probed the layer under shellImpl directly: spawn the resolved preferred shell running sleep 30 through CrossSpawnSpawner with the same options the shell tool uses (stdin: "ignore", forceKillAfter: "3 seconds"), start draining the merged output, then interrupt the fiber and let the scope release — ten iterations, 10/10 clean in 2–7ms. So interrupting a drain and killing the child does not hang at that level on Linux, and the close-versus-exit deviation recorded on #40 is not the mechanism here.
That leaves the cancel path above the spawner: prompt.cancel, shellImpl's abort handling, or the loop/queue coordination.
Why it matters beyond the test
If the cancel genuinely fails to end the shell about a quarter of the time, that is the product behaviour behind the test, not a test artifact: a user who aborts a long-running shell command gets a session that stays busy. The alternative — the cancel works and one of the fibers is not resolved — is a coordination bug in the same area. Both are worth more than a flaky test, and the current shape of the test cannot tell them apart.
Related but distinct: #40 is the Windows the shell never exited failure, whose own waits are bounded and named. Same family of question — a shell that ends and a wait that does not — on different platforms and, so far, with no shared mechanism established.
Found by accident, and it is not a Windows-runner story: it reproduces on a four-core Linux box in this container, with an unmodified checkout.
The measurement
packages/opencode/test/session/prompt.test.ts:2011, run on its own (bun test test/session/prompt.test.ts -t "cancel interrupts loop queued behind shell" --timeout 30000), ondevat37a3827c51with no local changes:One failure in four, and it is the test's own 30s ceiling rather than an assertion. Two neighbouring variants of an unrelated log line gave 0 failures in 4 and 2 in 3 — all small samples of the same flake, which is how it was noticed: I attributed a 30s hang to a line I had just moved, and the control run without that line hangs too. That attribution is retracted on #40.
What it is waiting for
The command is
sleep 30and the ceiling is 30s, so a cancel that does not actually end the shell is indistinguishable from the child simply running to completion. Since a failing assertion on the loop would report as an assertion, the timeout means one of the twoFiber.awaits never returned — and the test cannot say which, because the last one is unbounded. #28 gave the two sibling tests in this file named, bounded waits for exactly this reason; this one did not get the same treatment. Doing that is the cheap first step and would make the next occurrence say something.The spawner is not where it is stuck
Worth recording so nobody repeats it. I probed the layer under
shellImpldirectly: spawn the resolved preferred shell runningsleep 30throughCrossSpawnSpawnerwith the same options the shell tool uses (stdin: "ignore",forceKillAfter: "3 seconds"), start draining the merged output, then interrupt the fiber and let the scope release — ten iterations, 10/10 clean in 2–7ms. So interrupting a drain and killing the child does not hang at that level on Linux, and theclose-versus-exitdeviation recorded on #40 is not the mechanism here.That leaves the cancel path above the spawner:
prompt.cancel,shellImpl's abort handling, or the loop/queue coordination.Why it matters beyond the test
If the cancel genuinely fails to end the shell about a quarter of the time, that is the product behaviour behind the test, not a test artifact: a user who aborts a long-running shell command gets a session that stays busy. The alternative — the cancel works and one of the fibers is not resolved — is a coordination bug in the same area. Both are worth more than a flaky test, and the current shape of the test cannot tell them apart.
Related but distinct: #40 is the Windows
the shell never exitedfailure, whose own waits are bounded and named. Same family of question — a shell that ends and a wait that does not — on different platforms and, so far, with no shared mechanism established.