Repository navigation
Refuse webhooks the agent cannot deliver, and name the timeout in the log - #147
Merged
Merged
Conversation
… log Three faults in the webhook path, all of which turn a delayed handler into silently lost messages. The dispatch queue holds 100 invocations and Trigger sent to it with a blocking send. Trigger runs on the inbound HTTP goroutine, so webhook 101 hung that request with no timeout while the goroutines piled up. Make the send non-blocking and return ErrDispatchQueueFull instead. A refused invocation is cancelled, so it completes rather than sitting until its own timeout expires. The webhook handler answered 2xx as soon as it queued the work. When nothing was draining the queue the sender was told the webhook had been received, the invocation timed out unseen, and the message was gone. Return 503 with Retry-After on a full queue so the sender can retry. This is what Paychex reported: "axon is responding with a 2xx response telling us the webhook is being successfully received, but then it never is making it to our code." Both timeout paths report code "timeout" with no message, and the log line rendered the message alone, so the operator saw "invocation error: " and nothing else. Render the code when there is no message, and give the agent's own timer a message saying what it waited for. The handlerManager data races that -race reports here are pre-existing and unrelated, as is TestGRPCServer_ClientAutoClose, which fails on main. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
aszarama
previously approved these changes
Sep 14, 2026
keithfz
enabled auto-merge (squash)
September 14, 2026 17:28
The test registered the handler with the shared fixture option, which triggers on a 1ms interval, then started the handler. Those scheduled invocations went into the same dispatch queue that the test was filling, so the queue could reach its depth before the fill loop finished and the loop's own Trigger was refused. It failed 58 times in 100 local runs. Register a WEBHOOK option instead: the handler is active, but only the test puts work in the queue. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
aszarama
approved these changes
Sep 14, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three faults in the webhook path, all of which turn a delayed handler into silently lost messages. From CD-596.
Companion to #146, which removes the cause on the Python SDK side. This PR makes the agent behave correctly when a handler falls behind for any reason.
1. A full queue hung the inbound request
The dispatch queue holds 100 invocations, and
Triggerused a blocking send:Triggerruns on the inbound HTTP goroutine, so webhook 101 hung that request with no timeout, and goroutines piled up behind it. I hit this while writing a test — it never returned.The send is now non-blocking and returns
ErrDispatchQueueFull. A refused invocation is cancelled, so it completes immediately instead of sitting until its own timeout expires.2. The agent promised delivery it could not keep
The webhook handler returned 2xx as soon as it queued the work. When nothing was draining the queue, the sender was told the webhook had been received, the invocation timed out unseen, and the message was gone.
It now returns 503 with
Retry-Afteron a full queue, so the sender can retry. This is what Paychex described:3. The timeout never named itself
Both timeout paths report code
timeoutwith no message — the agent's own timer insendInvocations, and the SDK ataxon_agent.go:439. The log line rendered only the message:So the operator saw
"error":"invocation error: "with nothing after it, which is exactly what the ticket flagged as unexplained. It now falls back to the code, and the agent's timer carries a message saying what it waited for.Tests
TestTriggerRejectsWhenQueueIsFull— fills the queue, asserts the nextTriggerreturns rather than blocks, and that the refused invocation is completed. Fails against the blocking send (verified).TestHandleWebhookRejectsWhenQueueIsFull— 100 webhooks get 200, the next gets 503 and does not hang.TestDescribeError— covers code-only, message-only, both, and neither.Pre-existing failures, not from this change
Both reproduce on
main:TestGRPCServer_ClientAutoClosefails.go test -race ./server/handler/reports data races onhandlerManager's unsynchronised maps (14 warnings onmain). Worth its own ticket — a concurrent map write aborts the process, which restarts the container.🤖 Generated with Claude Code