feat(plugins): send a cancel notification to the plugin on call timeout - #842
Conversation
When a plugin call times out the host only dropped its own pending entry, so the plugin kept working and, for SQL drivers, the statement kept running on the server (#832). The management task now also writes a JSON-RPC notification to the plugin's stdin: {"jsonrpc":"2.0","method":"cancel","params":{"id":<request id>}} It is only sent while the request is still pending, so a response that raced the timeout does not trigger it, and never when the timeout is disabled. The contract is documented in PLUGIN_GUIDE.md.
| let lines: Vec<Value> = std::fs::read_to_string(log) | ||
| .unwrap_or_default() | ||
| .lines() | ||
| .map(|line| serde_json::from_str(line).unwrap()) |
There was a problem hiding this comment.
SUGGESTION: Fragile parse of a log file that is still being written
wait_for_lines polls stdin.log while the spawned cat is appending to it, so a poll can observe a partially written line (the last, not-yet-terminated chunk). serde_json::from_str(line).unwrap() then panics and the test fails intermittently instead of just seeing fewer lines. Filtering unparsable lines makes the helper tolerant of the concurrent writer:
| .map(|line| serde_json::from_str(line).unwrap()) | |
| .filter_map(|line| serde_json::from_str(line).ok()) |
Reply with @kilocode-bot fix it to have Kilo Code address this issue.
|
|
||
| ### Cancel Notification (Optional) | ||
|
|
||
| When a call exceeds the configured plugin call timeout, Tabularis stops waiting and reports the error to the user. Right after that it writes a JSON-RPC **notification** (no top-level `id`) naming the abandoned request: |
There was a problem hiding this comment.
SUGGESTION: Document that the cancel is only sent when the request is still pending
The host only writes the notification while the entry is still pending (if pending_requests.remove(&id).is_some() in driver.rs), so a response that raced the timeout produces no cancel. The wording "Right after that it writes" reads as unconditional; plugin authors could end up relying on always receiving a cancel. Worth stating that no cancel arrives for a request that already completed.
Reply with @kilocode-bot fix it to have Kilo Code address this issue.
Code Review SummaryStatus: 2 Issues Found | Recommendation: Address before merge Overview
Issue Details (click to expand)SUGGESTION
Files Reviewed (5 files)
Fix these issues in Kilo Cloud Reviewed by free · Input: 0 · Output: 0 · Cached: 0 |
|
Reviewed the
|
|
Merging it |
Stacked on #833 (base is
feat/configurable-plugin-call-timeout). Retarget tomainonce #833 is merged.Host side of #832. When a plugin call timed out, the host only dropped its own pending entry, so the plugin kept working and a Postgres statement kept running on the server (a
DELETEstill took effect after the user saw an error). The plugin side already ships inpostgresql-plugin1.0.0-rc.6 (TabularisDB/tabularis-postgresql-plugin#127), using the envelope agreed in the issue.What changed
plugins/rpc.rs: newJsonRpcNotificationtype andcancel_notification_line(id), which builds the line written to the plugin:{"jsonrpc":"2.0","method":"cancel","params":{"id":<request id>}}id, so it is a real JSON-RPC notification and the plugin must not reply.plugins/driver.rs: the management task already handledPluginCommand::Cancel(id)by removing the pending entry. It now also writes the cancel notification to stdin, but only if the entry was still pending: a response that raced the timeout does not trigger a cancel. With the timeout disabled (0, from feat(plugins): configurable plugin call timeout with per-plugin override #833) the timeout branch is never reached, so nothing is sent.plugins/PLUGIN_GUIDE.md: new "Cancel Notification (Optional)" section (do not reply, ignore unknown ids, keep reading stdin while a request runs).Plugins that do not handle
cancelkeep working: if they answer the notification anyway (e.g. with "method not found"), that line matches no pending request and is dropped, at worst with oneFailed to parse plugin responselog line.Verification
rpc_tests.rs: the line is a single newline-terminated notification, has no top-levelid, and keepsu64::MAXexact.cancel_tests.rs(unix, real child process via/bin/shrecording stdin): a timed-out call is followed by a cancel carrying the same id; a call with no timeout never sends one.cargo test --lib plugins::passes (147 tests), no clippy warnings in the touched files.SELECT pg_sleep(60)and a 3s timeout:activeduring the call, host error at 3.0s, no backend left 500ms later.activeafter the timeout.