What happens
Run HttpApi exerciser gates in the unit (linux) job stops producing output in the effect mode and never finishes. Twice so far, at the same point: right after the header that reports route selection, before the first scenario line.
1. dev at 719825686f (run 36043866195):
19:00:58 mode=effect selected=208 scenarioTimeout=30s effectRoutes=188 missing=0 extra=0
00:49:21 ##[error]The operation was canceled.
Five hours and forty-nine minutes of silence.
2. PR #21 at cb683a082d (run 36261980238):
18:28:04 mode=effect selected=208 scenarioTimeout=30s effectRoutes=188 missing=0 extra=0
19:19:28 ##[error]The operation was canceled.
Fifty-one minutes, and only because I cancelled it.
In both, the earlier coverage and auth modes had passed seconds before, and the job's cleanup terminated two orphan bun processes. On the runs where it works, the whole step takes about 2m31s (for instance dev at a697ccf091, half an hour before occurrence 2 and on the same code).
Each scenario already has its own 30s timeout, so this is not one slow scenario: the loop is stalling before it reports the first one.
What the step does about it: nothing
Only Run unit tests carries a timeout-minutes (20 on Linux, 50 on Windows). The exerciser step has none, so it inherits the job default of six hours. That is why occurrence 1 burned a runner most of a night and occurrence 2 sat on a PR until somebody noticed, instead of failing in twenty minutes and telling us which mode.
What I could and could not reproduce
Run alone, each of the three modes passes 208/208 and exits 0, in about a minute each. The hang does not reproduce here that way.
What did reproduce is narrower than the CI symptom, and I am recording it as a lead rather than a diagnosis. Two exercisers running at the same time — my own doing, two overlapping local runs — each printed summary pass=208 fail=0 skip=0 and then never exited; a 600s timeout killed both. That stall is at a different point than CI's (after the scenarios rather than before the first one), so it establishes that a second concurrent exerciser can hang, not that this is what CI hits. The two orphan bun processes at cleanup in both hung jobs are consistent with a previous mode's process still being alive when the next one starts, but nothing observed so far establishes that either.
Also, separately
A run leaves an untracked packages/opencode/config.json behind in the package root, containing only its $schema line. Small, but it is the same kind of thing: the exerciser writes and opens more than it cleans up.
What would help
- Bound the step with a
timeout-minutes so a stall there fails with a log instead of parking the job for hours. This alone turns the next occurrence into evidence.
- Make the stall observable: the mode prints its header and then nothing, so there is no way to tell what it is waiting on from the log we get.
- Establish whether a mode's process outlives its mode, since that is what the orphan pairs and the concurrent-run stall both point at.
What happens
Run HttpApi exerciser gatesin theunit (linux)job stops producing output in theeffectmode and never finishes. Twice so far, at the same point: right after the header that reports route selection, before the first scenario line.1.
devat719825686f(run 36043866195):Five hours and forty-nine minutes of silence.
2. PR #21 at
cb683a082d(run 36261980238):Fifty-one minutes, and only because I cancelled it.
In both, the earlier
coverageandauthmodes had passed seconds before, and the job's cleanup terminated two orphanbunprocesses. On the runs where it works, the whole step takes about 2m31s (for instancedevata697ccf091, half an hour before occurrence 2 and on the same code).Each scenario already has its own 30s timeout, so this is not one slow scenario: the loop is stalling before it reports the first one.
What the step does about it: nothing
Only
Run unit testscarries atimeout-minutes(20 on Linux, 50 on Windows). The exerciser step has none, so it inherits the job default of six hours. That is why occurrence 1 burned a runner most of a night and occurrence 2 sat on a PR until somebody noticed, instead of failing in twenty minutes and telling us which mode.What I could and could not reproduce
Run alone, each of the three modes passes 208/208 and exits 0, in about a minute each. The hang does not reproduce here that way.
What did reproduce is narrower than the CI symptom, and I am recording it as a lead rather than a diagnosis. Two exercisers running at the same time — my own doing, two overlapping local runs — each printed
summary pass=208 fail=0 skip=0and then never exited; a 600s timeout killed both. That stall is at a different point than CI's (after the scenarios rather than before the first one), so it establishes that a second concurrent exerciser can hang, not that this is what CI hits. The two orphanbunprocesses at cleanup in both hung jobs are consistent with a previous mode's process still being alive when the next one starts, but nothing observed so far establishes that either.Also, separately
A run leaves an untracked
packages/opencode/config.jsonbehind in the package root, containing only its$schemaline. Small, but it is the same kind of thing: the exerciser writes and opens more than it cleans up.What would help
timeout-minutesso a stall there fails with a log instead of parking the job for hours. This alone turns the next occurrence into evidence.