Retry transient API errors and ease load on the nightly compositional shard - #164
Merged
Merged
Conversation
… shard The nightly shard was failing intermittently under -n 20: some failures were Playwright timeouts that exhausted their retries because the pool was genuinely overloaded, but others were failures the existing retry never covered — 502 Bad Gateway responses from the Table API, and a "record was not created" check that gave up after ~7.5s instead of the 30s budget used elsewhere for browser-side waits. Retry transient 502/503/504 and connection errors at the request layer, widen the record-creation poll to match SNOW_BROWSER_TIMEOUT, and drop the shard to -n 10 to reduce concurrent load on the shared pool. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The re-run of the nightly shard on this branch dropped from 9 failures to
1: iframe.get_by_label("All").click() in the change-request creation flow
hit a Playwright strict-mode violation because "All" substring-matches
both the intended category-filter link and the list view's unrelated
"Select All" checkbox. Disambiguate with .first, matching the identical
pattern already used one line below for "Normal". Verified against the
live pool by reproducing the exact failing case (seed 271, level 3,
NavigateAndCreateChangeRequestTask) before and after the fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The nightly shard (
instance_pool_ci.yml→test-nightly-shard) has been failing intermittently under-n 20(e.g. run 36144671490, 9 failures). Digging into that run's logs, the 9 failures split into two categories:TimeoutErrors that exhausted the existing 5-attempt retry — genuine sustained load on the pool, not fixable in code.502 Bad Gatewayresponses from the Table API (table_api_callhad no retry on the HTTP request itself, only on "record not yet visible"), and 1ValueError: The record was not createdfrom a poll inform.pythat gave up after ~7.5s instead of the 30s budget (SNOW_BROWSER_TIMEOUT) used elsewhere for browser-side waits.Changes:
api/utils.py: added_request_with_retry, retrying transient502/503/504and connection errors with exponential backoff (5 attempts), used bytable_api_call,table_column_info, anddb_delete_from_table.tasks/form.py: widened the "record was not created" poll to the same 30s (SNOW_BROWSER_TIMEOUT) budget instead of a hardcoded ~7.5s.instance_pool_ci.yml: dropped the nightly shard from-n 20to-n 10to reduce concurrent load on the shared pool (the genuinely load-caused timeouts).Update: dispatched the nightly shard on this branch (run 36446668915) — down to 1 failure from 9, confirming the retry/backoff and lower concurrency worked. The one remaining failure was a genuine, load-independent bug:
iframe.get_by_label("All").click()in the change-request creation cheat hits a Playwright strict-mode violation because"All"substring-matches both the intended category-filter link and an unrelated "Select All" checkbox. Fixed with.first, matching the identical pattern already used one line below for"Normal"(and precedented in #162 for the same class of bug). Verified by reproducing the exact failing case (seed 271, level 3,NavigateAndCreateChangeRequestTask) against the live pool before and after the fix.Test plan
black==24.2.0(CI's pinned version) unchanged._request_with_retryretries on502and backs off, but fails fast on real errors like404(not left in the repo — ad hoc verification).-m 'not slow and not pricy and not pool_health') passes against the live pool: 66 passed.🤖 Generated with Claude Code