Skip to content

The HttpApi exerciser gate hangs in effect mode, and the step it runs in has no timeout #22

Description

@Edo771977

What happens

Run HttpApi exerciser gates in the unit (linux) job stops producing output in the effect mode and never finishes. Twice so far, at the same point: right after the header that reports route selection, before the first scenario line.

1. dev at 719825686f (run 36043866195):

19:00:58  mode=effect selected=208 scenarioTimeout=30s effectRoutes=188 missing=0 extra=0
00:49:21  ##[error]The operation was canceled.

Five hours and forty-nine minutes of silence.

2. PR #21 at cb683a082d (run 36261980238):

18:28:04  mode=effect selected=208 scenarioTimeout=30s effectRoutes=188 missing=0 extra=0
19:19:28  ##[error]The operation was canceled.

Fifty-one minutes, and only because I cancelled it.

In both, the earlier coverage and auth modes had passed seconds before, and the job's cleanup terminated two orphan bun processes. On the runs where it works, the whole step takes about 2m31s (for instance dev at a697ccf091, half an hour before occurrence 2 and on the same code).

Each scenario already has its own 30s timeout, so this is not one slow scenario: the loop is stalling before it reports the first one.

What the step does about it: nothing

Only Run unit tests carries a timeout-minutes (20 on Linux, 50 on Windows). The exerciser step has none, so it inherits the job default of six hours. That is why occurrence 1 burned a runner most of a night and occurrence 2 sat on a PR until somebody noticed, instead of failing in twenty minutes and telling us which mode.

What I could and could not reproduce

Run alone, each of the three modes passes 208/208 and exits 0, in about a minute each. The hang does not reproduce here that way.

What did reproduce is narrower than the CI symptom, and I am recording it as a lead rather than a diagnosis. Two exercisers running at the same time — my own doing, two overlapping local runs — each printed summary pass=208 fail=0 skip=0 and then never exited; a 600s timeout killed both. That stall is at a different point than CI's (after the scenarios rather than before the first one), so it establishes that a second concurrent exerciser can hang, not that this is what CI hits. The two orphan bun processes at cleanup in both hung jobs are consistent with a previous mode's process still being alive when the next one starts, but nothing observed so far establishes that either.

Also, separately

A run leaves an untracked packages/opencode/config.json behind in the package root, containing only its $schema line. Small, but it is the same kind of thing: the exerciser writes and opens more than it cleans up.

What would help

  • Bound the step with a timeout-minutes so a stall there fails with a log instead of parking the job for hours. This alone turns the next occurrence into evidence.
  • Make the stall observable: the mode prints its header and then nothing, so there is no way to tell what it is waiting on from the log we get.
  • Establish whether a mode's process outlives its mode, since that is what the orphan pairs and the concurrent-run stall both point at.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions