Skip to content

Retry transient API errors and ease load on the nightly compositional shard - #164

Merged
aldro61 merged 2 commits into
mainfrom
fix/nightly-shard-transient-errors
Sep 28, 2026
Merged

aldro61 merged 2 commits into
mainfrom
fix/nightly-shard-transient-errors

Conversation

@aldro61

@aldro61 aldro61 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The nightly shard (instance_pool_ci.yml → test-nightly-shard) has been failing intermittently under -n 20 (e.g. run 36144671490, 9 failures). Digging into that run's logs, the 9 failures split into two categories:

  • 6 were Playwright TimeoutErrors that exhausted the existing 5-attempt retry — genuine sustained load on the pool, not fixable in code.
  • 3 were failures the existing retry never covered at all: 2 502 Bad Gateway responses from the Table API (table_api_call had no retry on the HTTP request itself, only on "record not yet visible"), and 1 ValueError: The record was not created from a poll in form.py that gave up after ~7.5s instead of the 30s budget (SNOW_BROWSER_TIMEOUT) used elsewhere for browser-side waits.

Changes:

  • api/utils.py: added _request_with_retry, retrying transient 502/503/504 and connection errors with exponential backoff (5 attempts), used by table_api_call, table_column_info, and db_delete_from_table.
  • tasks/form.py: widened the "record was not created" poll to the same 30s (SNOW_BROWSER_TIMEOUT) budget instead of a hardcoded ~7.5s.
  • instance_pool_ci.yml: dropped the nightly shard from -n 20 to -n 10 to reduce concurrent load on the shared pool (the genuinely load-caused timeouts).

Update: dispatched the nightly shard on this branch (run 36446668915) — down to 1 failure from 9, confirming the retry/backoff and lower concurrency worked. The one remaining failure was a genuine, load-independent bug: iframe.get_by_label("All").click() in the change-request creation cheat hits a Playwright strict-mode violation because "All" substring-matches both the intended category-filter link and an unrelated "Select All" checkbox. Fixed with .first, matching the identical pattern already used one line below for "Normal" (and precedented in #162 for the same class of bug). Verified by reproducing the exact failing case (seed 271, level 3, NavigateAndCreateChangeRequestTask) against the live pool before and after the fix.

Test plan

  • All modified files pass black==24.2.0 (CI's pinned version) unchanged.
  • Isolated mocked tests confirm _request_with_retry retries on 502 and backs off, but fails fast on real errors like 404 (not left in the repo — ad hoc verification).
  • Fast per-PR suite (-m 'not slow and not pricy and not pool_health') passes against the live pool: 66 passed.
  • Nightly shard re-run on this branch: 1 failed → 0 failed after the locator fix (1 failed, 98 passed → confirmed passing in isolation with the fix; full shard re-run pending).

🤖 Generated with Claude Code

aldro61 and others added 2 commits September 28, 2026 11:46
… shard

The nightly shard was failing intermittently under -n 20: some failures
were Playwright timeouts that exhausted their retries because the pool
was genuinely overloaded, but others were failures the existing retry
never covered — 502 Bad Gateway responses from the Table API, and a
"record was not created" check that gave up after ~7.5s instead of the
30s budget used elsewhere for browser-side waits. Retry transient
502/503/504 and connection errors at the request layer, widen the
record-creation poll to match SNOW_BROWSER_TIMEOUT, and drop the shard
to -n 10 to reduce concurrent load on the shared pool.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The re-run of the nightly shard on this branch dropped from 9 failures to
1: iframe.get_by_label("All").click() in the change-request creation flow
hit a Playwright strict-mode violation because "All" substring-matches
both the intended category-filter link and the list view's unrelated
"Select All" checkbox. Disambiguate with .first, matching the identical
pattern already used one line below for "Normal". Verified against the
live pool by reproducing the exact failing case (seed 271, level 3,
NavigateAndCreateChangeRequestTask) before and after the fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@aldro61
aldro61 merged commit e885130 into main Sep 28, 2026
14 checks passed
@aldro61
aldro61 deleted the fix/nightly-shard-transient-errors branch September 28, 2026 19:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant