Skip to content

A PTY exit goes unobserved on Linux, and something stalls a shell loop test on Windows #11

Description

@Edo771977

Rewritten on 2026-09-27. This issue started as "dev is red at 0bd1b876, and the Windows unit job is the only job failing", and over five occurrences it accumulated four separate things. Two are fixed, one was diagnosed and fixed, and the earlier reading of the fourth turned out to be wrong. The comments below are the record and are left as they were written; this description is what is actually still open.

Still open

1. A PTY exit goes unobserved (Linux)

From the CI on #10:

test/server/httpapi-v2-pty.test.ts:93
(fail) v2 pty HttpApi > serves location-wrapped PTY routes and retains exited sessions [21256.87ms]
-   "status": "exited",  "exitCode": 4
+   "status": "running", "pid": 6907, …

The test polls every 50ms against a 20 second deadline, and sh -c "exit 4" was still reported as running when it expired. The Linux job ran 557s against ~450s on dev — 24% slower, which does not account for twenty seconds, so this is not the slow-runner family.

It did not reproduce on a re-run of the same commit and passes 6/6 locally, so it is intermittent. That is not a reason to file it away: a run that says a process is still running twenty seconds after it exited is either a lost exit event or a poll that never re-reads the real state, and both are worth finding. Nothing has been done about this one. It is the oldest thing in this issue and the only one still untouched.

(For anyone reproducing locally: a different test in that file, applies plugin shell environment before forced PTY values, fails in a sandbox that cannot run bun install. That is not this bug.)

2. Something stalls one of the two shell loop tests on Windows — and it is not the ceiling

loop waits while shell runs and starts after shell exits timed out at 30002.99ms on 968ba5e490 (job).

This issue used to read that as the timing drift crossing the 30s ceiling #10 had given it. That reading does not survive measurement. On Linux, with bun's junit reporter over the whole file:

test body time ceiling
loop waits while shell runs and starts after shell exits 1.83s 30s
shell completion resumes queued loop callers 2.53s 30s

12 to 16 times the headroom they need. The Windows unit job runs about two and a half times slower, which would put them at 4.6s and 6.3s. Nothing about a slow machine turns 1.8s into 30s while the same machine finishes 3700 other tests: something waited. Raising the ceiling again would have been choosing a number for a phenomenon nobody had identified.

What was missing was knowing what it waited for. Both tests forked the shell and the loop and awaited them with bare Fiber.await, so whichever side stopped, the test died with this test timed out after 30000ms, naming neither. Two occurrences — at 10s under the old ceiling, at 30s under the current one — produced no information about the cause for exactly that reason.

#28 (16bea30769) bounds and names both waits. It does not fix this: if something stalls on Windows the test still fails, but it now fails saying the shell never exited or a queued loop caller never resumed after the shell exited. So this item is open as "one of those two waits stalls on Windows, cause unknown", and the next occurrence will say which one. The ceiling stays at 30s deliberately; the measurement says it is not the constraint.

Fixed, kept here for the record

what where it went
Skill refresh kept the cached copy on Windows. Diagnosed at last as EPERM on rename(root, backup) — the swap refused while a handle was held inside the directory, the download discarded, the stale copy returned as if current #25 → #26 (7361fc46ef): the swap is retried, bounded, in both runtimes, and the rollback that could lose the skill outright is covered
Logs from these suites were invisible, so the absence of failed to refresh skill in a Windows log was read as evidence that the rename had not failed — it was evidence of nothing #12 → #13 (719825686f): a failing test replays its captured console output. The very first recurrence after it handed over the syscall, both paths and the errno, which is what made the fix above possible
Three tests carried ceilings set against a fast runner: the Azure Node compatibility test at 30s, and the two shell loop tests at 10s #10 (062b843039): 120s for Azure, 30s for the two shell tests
A fourth test, same class: InstanceStore.provide runs InstanceBootstrap before effect crossing the suite-wide 30s #20 → #21 (2add03aeed): its file's tests carry their own ceiling, with the measurement in the comment — 14.7s for the first test, 0.4s for each one after, because whichever runs first pays the process's first real InstanceBootstrap

What this issue taught, twice

Both times the diagnosis was wrong in the same direction: a failure that looked like slowness was read as slowness, because the only thing the failure said was that time had run out. The skill refresh needed logs replayed (#13) before it could say EPERM; the shell tests needed their waits named (#28) before they can say which side stopped. Neither was fixed by choosing a bigger number.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions