Skip to content

Refuse webhooks the agent cannot deliver, and name the timeout in the log - #147

Merged
keithfz merged 4 commits into
mainfrom
keithzeto/cd-596-agent-webhook-backpressure
Sep 14, 2026
Merged

keithfz merged 4 commits into
mainfrom
keithzeto/cd-596-agent-webhook-backpressure

Conversation

@keithfz

@keithfz keithfz commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Three faults in the webhook path, all of which turn a delayed handler into silently lost messages. From CD-596.

Companion to #146, which removes the cause on the Python SDK side. This PR makes the agent behave correctly when a handler falls behind for any reason.

1. A full queue hung the inbound request

The dispatch queue holds 100 invocations, and Trigger used a blocking send:

queue <- message

Trigger runs on the inbound HTTP goroutine, so webhook 101 hung that request with no timeout, and goroutines piled up behind it. I hit this while writing a test — it never returned.

The send is now non-blocking and returns ErrDispatchQueueFull. A refused invocation is cancelled, so it completes immediately instead of sitting until its own timeout expires.

2. The agent promised delivery it could not keep

The webhook handler returned 2xx as soon as it queued the work. When nothing was draining the queue, the sender was told the webhook had been received, the invocation timed out unseen, and the message was gone.

It now returns 503 with Retry-After on a full queue, so the sender can retry. This is what Paychex described:

axon is responding with a 2xx response telling us the webhook is being successfully received, but then it never is making it to our code for the webhook for processing because of the failures. This is making lost messages.

3. The timeout never named itself

Both timeout paths report code timeout with no message — the agent's own timer in sendInvocations, and the SDK at axon_agent.go:439. The log line rendered only the message:

fmt.Errorf("invocation error: %s", req.GetError().GetMessage())

So the operator saw "error":"invocation error: " with nothing after it, which is exactly what the ticket flagged as unexplained. It now falls back to the code, and the agent's timer carries a message saying what it waited for.

Tests

  • TestTriggerRejectsWhenQueueIsFull — fills the queue, asserts the next Trigger returns rather than blocks, and that the refused invocation is completed. Fails against the blocking send (verified).
  • TestHandleWebhookRejectsWhenQueueIsFull — 100 webhooks get 200, the next gets 503 and does not hang.
  • TestDescribeError — covers code-only, message-only, both, and neither.

Pre-existing failures, not from this change

Both reproduce on main:

  • TestGRPCServer_ClientAutoClose fails.
  • go test -race ./server/handler/ reports data races on handlerManager's unsynchronised maps (14 warnings on main). Worth its own ticket — a concurrent map write aborts the process, which restarts the container.

🤖 Generated with Claude Code

… log

Three faults in the webhook path, all of which turn a delayed handler
into silently lost messages.

The dispatch queue holds 100 invocations and Trigger sent to it with a
blocking send. Trigger runs on the inbound HTTP goroutine, so webhook 101
hung that request with no timeout while the goroutines piled up. Make the
send non-blocking and return ErrDispatchQueueFull instead. A refused
invocation is cancelled, so it completes rather than sitting until its
own timeout expires.

The webhook handler answered 2xx as soon as it queued the work. When
nothing was draining the queue the sender was told the webhook had been
received, the invocation timed out unseen, and the message was gone.
Return 503 with Retry-After on a full queue so the sender can retry.
This is what Paychex reported: "axon is responding with a 2xx response
telling us the webhook is being successfully received, but then it never
is making it to our code."

Both timeout paths report code "timeout" with no message, and the log
line rendered the message alone, so the operator saw "invocation error: "
and nothing else. Render the code when there is no message, and give the
agent's own timer a message saying what it waited for.

The handlerManager data races that -race reports here are pre-existing
and unrelated, as is TestGRPCServer_ClientAutoClose, which fails on main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@keithfz
keithfz requested review from a team, aszarama and shawnburke September 11, 2026 20:20
aszarama
aszarama previously approved these changes Sep 14, 2026
@keithfz
keithfz enabled auto-merge (squash) September 14, 2026 17:28
The test registered the handler with the shared fixture option, which
triggers on a 1ms interval, then started the handler. Those scheduled
invocations went into the same dispatch queue that the test was filling,
so the queue could reach its depth before the fill loop finished and the
loop's own Trigger was refused. It failed 58 times in 100 local runs.

Register a WEBHOOK option instead: the handler is active, but only the
test puts work in the queue.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@keithfz
keithfz merged commit e125b9a into main Sep 14, 2026
21 checks passed
@keithfz
keithfz deleted the keithzeto/cd-596-agent-webhook-backpressure branch September 14, 2026 20:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants