Conversation
execute_query_batch already runs its statements on one pooled connection,
so BEGIN ... COMMIT inside a script works. The workflow the transaction
exists for did not: run BEGIN, inspect, run the changes, verify, then
COMMIT, each as its own run. Between runs the connection went back to the
pool, so the next run could land on a different one - the verify step saw
pre-transaction data and COMMIT reported no transaction in progress while
the real one stayed open elsewhere.
The pool recycles with RecyclingMethod::Fast, which resets nothing, so
that stranded connection also went back still inside the transaction,
holding its locks, and the next borrower ran inside it.
- transaction_effect() classifies a statement's effect on the surrounding
transaction from its leading keywords only, so a BEGIN in a string
literal or a PL/pgSQL block body is not mistaken for transaction
control, and ROLLBACK TO SAVEPOINT correctly leaves it open.
- session: registry of connections pinned per host session, a ROLLBACK
before any pinned connection returns to the pool, and an idle sweep so
an abandoned session cannot hold locks indefinitely.
- execute_query_batch honours session_id and answers with
{ results, in_transaction } when one is given, keeping the bare array
for hosts that do not send one.
- release_session RPC method; shutdown releases every pinned session.
- search_path is not re-applied to a reused connection: it would run
inside the open transaction and change what the rest of it sees.
Running statements one at a time is exactly how a transaction is driven -
BEGIN, the changes, a verifying SELECT, COMMIT - and the host sends each
of those through execute_query, not execute_query_batch. Without session
handling there, every run took a fresh pooled connection: the UPDATE
auto-committed on its own connection and the later COMMIT/ROLLBACK hit a
connection with no transaction, where PostgreSQL only warns and reports
success.
- Extract run_batch_in_session() from execute_query_batch so a single
statement and a batch take the same session path.
- execute_query honours session_id and answers with
{ result, in_transaction } when one is given, keeping the bare
QueryResult for hosts that do not send one.
|
Thanks for this, @egertaia — this is a genuinely well-scoped fix for a real correctness/lock-leak bug, and it's unusually well self-documented for a PR of this size. The write-up connecting One non-blocking note on hardening, plus a live-DB pass that turned up something I think does need fixing before this merges — not just a nice-to-have. 1. Idle-sweep is opportunistic, not scheduled (non-blocking)
// main.rs
async fn run_pool_cleanup(mut shutdown_rx: watch::Receiver<bool>) {
let mut timer = interval(POOL_CLEANUP_INTERVAL);
loop {
tokio::select! {
_ = timer.tick() => {
postgresql_plugin::client::cleanup_idle_pools();
postgresql_plugin::session::sweep_idle().await;
}
_ = shutdown_rx.changed() => break,
}
}
}with 2. Live-database coverage — draft tests, and a bug I'd call blockingYou mentioned in the PR description you'd welcome pointers on wiring this into
Ready-to-drop-in versions of all of these, adapted to fn setup_session_scratch_table(plugin: &mut Plugin, params: &Value) {
plugin.call_ok("execute_query", json!({
"params": params,
"query": "CREATE TABLE IF NOT EXISTS live_db_session_scratch \
(id SERIAL PRIMARY KEY, value INTEGER)",
}));
plugin.call_ok("execute_query", json!({
"params": params,
"query": "TRUNCATE live_db_session_scratch RESTART IDENTITY",
}));
plugin.call_ok("execute_query", json!({
"params": params,
"query": "INSERT INTO live_db_session_scratch (value) VALUES (1)",
}));
}
#[test]
fn session_id_keeps_a_transaction_alive_across_separate_execute_query_calls() {
let mut plugin = Plugin::spawn();
let params = conn_params();
setup_session_scratch_table(&mut plugin, ¶ms);
let session = json!("live-db-session-1");
let r1 = plugin.call_ok("execute_query", json!({
"params": params, "session_id": session, "query": "BEGIN"
}));
assert_eq!(r1["in_transaction"], json!(true));
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session,
"query": "UPDATE live_db_session_scratch SET value = 99 WHERE id = 1"
}));
let r2 = plugin.call_ok("execute_query", json!({
"params": params, "session_id": session,
"query": "SELECT value FROM live_db_session_scratch WHERE id = 1"
}));
assert_eq!(
r2["result"]["rows"][0][0], json!(99),
"verifying SELECT on the same session must see the uncommitted UPDATE"
);
let r3 = plugin.call_ok("execute_query", json!({
"params": params, "session_id": session, "query": "COMMIT"
}));
assert_eq!(r3["in_transaction"], json!(false));
let r4 = plugin.call_ok("execute_query", json!({
"params": params,
"query": "SELECT value FROM live_db_session_scratch WHERE id = 1"
}));
assert_eq!(r4["rows"][0][0], json!(99), "committed value must be visible with no session_id");
}
#[test]
fn rollback_discards_changes_made_across_separate_calls() {
let mut plugin = Plugin::spawn();
let params = conn_params();
setup_session_scratch_table(&mut plugin, ¶ms);
let session = json!("live-db-session-2");
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session, "query": "BEGIN"
}));
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session,
"query": "UPDATE live_db_session_scratch SET value = 777 WHERE id = 1"
}));
let rollback = plugin.call_ok("execute_query", json!({
"params": params, "session_id": session, "query": "ROLLBACK"
}));
assert_eq!(rollback["in_transaction"], json!(false));
let check = plugin.call_ok("execute_query", json!({
"params": params,
"query": "SELECT value FROM live_db_session_scratch WHERE id = 1"
}));
assert_eq!(check["rows"][0][0], json!(1), "ROLLBACK must discard the UPDATE");
}
#[test]
fn abandoned_transaction_without_session_id_does_not_leak_its_lock() {
let mut plugin = Plugin::spawn();
let params = conn_params();
setup_session_scratch_table(&mut plugin, ¶ms);
// A batch with no session_id that leaves a transaction open.
plugin.call_ok("execute_query_batch", json!({
"params": params,
"queries": json!([
"BEGIN",
"UPDATE live_db_session_scratch SET value = 5 WHERE id = 1"
]),
}));
// If the abandoned transaction leaked back into the pool with its
// lock still held, this NOWAIT query on the same row would error.
let r = plugin.call("execute_query", json!({
"params": params,
"query": "SELECT value FROM live_db_session_scratch WHERE id = 1 FOR UPDATE NOWAIT"
}));
assert!(
r.get("error").is_none(),
"unrelated query must not inherit the abandoned transaction's lock: {r:?}"
);
}
#[test]
fn release_session_rolls_back_an_open_transaction() {
let mut plugin = Plugin::spawn();
let params = conn_params();
setup_session_scratch_table(&mut plugin, ¶ms);
let session = json!("live-db-session-3");
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session, "query": "BEGIN"
}));
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session,
"query": "UPDATE live_db_session_scratch SET value = 333 WHERE id = 1"
}));
plugin.call_ok("release_session", json!({ "session_id": session }));
let r = plugin.call_ok("execute_query", json!({
"params": params,
"query": "SELECT value FROM live_db_session_scratch WHERE id = 1"
}));
assert_eq!(r["rows"][0][0], json!(1), "release_session must roll back the uncommitted UPDATE");
}
#[test]
fn two_sessions_stay_isolated_from_each_other() {
let mut plugin = Plugin::spawn();
let params = conn_params();
setup_session_scratch_table(&mut plugin, ¶ms);
plugin.call_ok("execute_query", json!({
"params": params,
"query": "INSERT INTO live_db_session_scratch (value) VALUES (0)"
}));
let session_a = json!("live-db-session-4a");
let session_b = json!("live-db-session-4b");
plugin.call_ok("execute_query", json!({ "params": params, "session_id": session_a, "query": "BEGIN" }));
plugin.call_ok("execute_query", json!({ "params": params, "session_id": session_b, "query": "BEGIN" }));
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session_a,
"query": "UPDATE live_db_session_scratch SET value = 111 WHERE id = 1"
}));
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session_b,
"query": "UPDATE live_db_session_scratch SET value = 222 WHERE id = 2"
}));
let ra = plugin.call_ok("execute_query", json!({
"params": params, "session_id": session_a,
"query": "SELECT value FROM live_db_session_scratch WHERE id = 1"
}));
assert_eq!(ra["result"]["rows"][0][0], json!(111));
plugin.call_ok("execute_query", json!({ "params": params, "session_id": session_a, "query": "ROLLBACK" }));
plugin.call_ok("execute_query", json!({ "params": params, "session_id": session_b, "query": "COMMIT" }));
let r1 = plugin.call_ok("execute_query", json!({
"params": params, "query": "SELECT value FROM live_db_session_scratch WHERE id = 1"
}));
assert_eq!(r1["rows"][0][0], json!(1), "session_a's rollback must discard its own write");
let r2 = plugin.call_ok("execute_query", json!({
"params": params, "query": "SELECT value FROM live_db_session_scratch WHERE id = 2"
}));
assert_eq!(r2["rows"][0][0], json!(222), "session_b's commit must persist");
}
#[test]
fn batch_and_single_statement_calls_share_the_same_session() {
let mut plugin = Plugin::spawn();
let params = conn_params();
setup_session_scratch_table(&mut plugin, ¶ms);
let session = json!("live-db-session-5");
let r1 = plugin.call_ok("execute_query_batch", json!({
"params": params, "session_id": session, "queries": json!(["BEGIN"])
}));
assert_eq!(r1["in_transaction"], json!(true));
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session,
"query": "UPDATE live_db_session_scratch SET value = 55 WHERE id = 1"
}));
let r2 = plugin.call_ok("execute_query", json!({
"params": params, "session_id": session,
"query": "SELECT value FROM live_db_session_scratch WHERE id = 1"
}));
assert_eq!(r2["result"]["rows"][0][0], json!(55));
let r3 = plugin.call_ok("execute_query_batch", json!({
"params": params, "session_id": session, "queries": json!(["COMMIT"])
}));
assert_eq!(r3["in_transaction"], json!(false));
}The 8th scenario found a bug that I think should block merge, not follow up later, because it undermines the exact thing this PR sets out to fix — stray pinned connections outliving the transaction that justified holding them.
#[test]
fn commit_that_fails_on_a_deferred_constraint_does_not_leave_a_stale_pin() {
let mut plugin = Plugin::spawn();
let params = conn_params();
plugin.call_ok("execute_query", json!({
"params": params, "query": "DROP TABLE IF EXISTS live_db_deferred_fk_scratch"
}));
plugin.call_ok("execute_query", json!({
"params": params,
"query": "CREATE TABLE live_db_deferred_fk_scratch (id INT PRIMARY KEY, other_id INT)"
}));
plugin.call_ok("execute_query", json!({
"params": params,
"query": "ALTER TABLE live_db_deferred_fk_scratch ADD CONSTRAINT fk_self \
FOREIGN KEY (other_id) REFERENCES live_db_deferred_fk_scratch(id) \
DEFERRABLE INITIALLY DEFERRED"
}));
let session = json!("live-db-session-commit-fails");
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session, "query": "BEGIN"
}));
// The FK check is deferred, so this INSERT succeeds; only COMMIT fails.
plugin.call_ok("execute_query", json!({
"params": params, "session_id": session,
"query": "INSERT INTO live_db_deferred_fk_scratch (id, other_id) VALUES (1, 999)"
}));
let commit = plugin.call("execute_query", json!({
"params": params, "session_id": session, "query": "COMMIT"
}));
assert!(commit.get("error").is_some(), "COMMIT violating a deferred FK must error");
// PostgreSQL implicitly ends the transaction attempt once COMMIT
// itself fails — a later ROLLBACK on the same backend just warns
// "no transaction in progress". So the session should NOT still
// report in_transaction: true here.
let after = plugin.call_ok("execute_query", json!({
"params": params, "session_id": session, "query": "SELECT 1 AS still_usable"
}));
assert_eq!(
after["in_transaction"], json!(false),
"a failed COMMIT already ended the transaction server-side; \
the session must not stay pinned believing one is still open"
);
}Ran this against the branch locally and it fails: I verified the underlying Postgres behavior directly with A Really appreciate the depth here — happy to take another pass as soon as you've had a chance to look at the |
Version suggestionBased on this PR's title (
This is informational only — no tag or release is created automatically yet. |
PostgreSQL rolls the transaction back when COMMIT fails on its own (e.g. a deferred constraint), so keeping the session pinned left a connection held with no transaction behind it. Also classify ABORT, PREPARE TRANSACTION, COMMIT/ROLLBACK PREPARED and ... AND CHAIN, the last of which previously returned a connection to the pool with a fresh transaction still open.
The idle check only ran inside take(), so an abandoned session kept its transaction's locks until some other session happened to run a query. Hang sweep_idle() off the existing 10-minute pool cleanup timer instead.
|
Thanks, both confirmed and fixed on the branch (09849ce, 14548f4).
Your deferred-FK test is in Idle sweep: not deliberate, it was an oversight. |
|
Thanks for the quick turnaround, @egertaia — both fixes are correct and the audit you did beyond just my one repro case is the right instinct. What I checkedRather than just reading the diff, I built the release binary off this branch and ran it against a real, freshly-initialized PostgreSQL 16:
No blockers from any of that. One new issue, found doing a deeper concurrency pass — I think this should block mergeThe plugin dispatches requests across Reproduction 1 — stale read. Session Reproduction 2 — silent data loss, worse. Same setup, but both calls open a transaction and write to different rows (so they don't block on Postgres row locks — this is a plugin-side session-map race, not a database lock issue). Call B (fast) does This is reachable in real usage, not just theoretical. I checked the host side (
So a user pressing the run shortcut twice quickly, or right-clicking "Execute Selection" while a previous statement on the same tab is still executing, sends two overlapping Suggested fix direction: serialize per- |
Two overlapping calls for one session both found nothing pinned, each took a fresh connection, and the later store() rolled back the other's transaction. Each session now has its own lock held for the whole take, run, pin sequence, so an overlapping call waits its turn on the same connection, and release_session waits for a run still in flight.
It opens a new transaction whatever the prior state was.
|
Thanks for the repro, confirmed and fixed in 7005c15. Each session now has its own lock. I went with waiting rather than returning an error: statements then run in the order the user sent them, which is what the editor expects. Also on the branch: 6e9857c makes a successful |
Also covers the sweep forgetting a slot nothing uses.
|
@egertaia — ran this through the same review process I use on my own work — eleven independent lenses (correctness, security, concurrency, error-handling, contract-consistency, docs-vs-code, edge-cases, test-coverage, validation, backward-compat, test-structure), each looking only through its own lens. That fan-out raised a batch of raw findings, but the ones below are the ones that survived actually running them against the real binary and the real function — not just re-reading the diff and reasoning about what "should" happen. That distinction mattered a lot here: the most serious one (#1) only became undeniable once I instrumented the live code and watched it happen; a plausible-sounding version of it could easily have been argued away on paper. Sorry to come back with more after saying this looked close; better now than after merge, and happy to help however's useful — pair on any of these, or I can just take a pass at one/all of the fixes myself if that's faster for you. 1.
|
…CTION ROLLBACK WORK/TRANSACTION TO SAVEPOINT was classified as closing, returning a connection to the pool mid-transaction, and only leading comments were skipped, so COMMIT -- and chain read as chaining and START /* x */ TRANSACTION was missed. transaction_effect now reads keywords with every comment skipped and drops the optional noise word.
release_all drained every slot before skipping busy ones, so a run still in flight pinned its connection into a slot nothing tracked, and it went back to the pool without a ROLLBACK. Busy slots now stay in the map, and a slot dropped while holding a connection closes it instead of returning it.
A failing statement came back as an RPC error, so the host never learned that a failed COMMIT had ended the transaction. With a session_id it is now a reply carrying the error and in_transaction.
|
Thanks, and thanks for driving #1 live, that trace made it unambiguous. All four are fixed on the branch, each with a test that fails on the previous code. 1. 2 and 3. Classifier (54b5b13). I took the robust route: 4.
|
|
@egertaia — thank you for the work on this, genuinely. This was a long back-and-forth and you tracked down and fixed every single thing that came up, including the two subtle ones (the busy-slot orphan and the classifier's comment/keyword handling) that needed a live repro to even pin down. The I re-verified all four fixes independently rather than just reading the diffs:
Full suite: This is ready to merge from my side. Nice work getting it across the line. |
|
Closing this fork-based PR in favor of a fresh PR opened from All of @egertaia's original commits are preserved with full attribution in the replacement branch. The replacement is fully verified: 83/83 cross-repo parity, 356 unit tests, 30 live-DB tests (including the deferred-constraint session test, now confirmed passing against a live PostgreSQL instance), clippy/fmt/markdownlint clean. No reflection on the work here — this is purely the fork-PR mechanics (the head branch can't be updated from the upstream side). The replacement PR will link here. |
Companion to TabularisDB/tabularis#801, which plumbs the session id through the host. Either can merge first: the wire change is backward compatible in both directions.
Problem
execute_query_batchacquires one pooled connection for the whole batch, soBEGIN … COMMITinside a single script works. The workflow the transaction exists for does not:Each run is its own RPC call (and a one-statement run arrives as
execute_query, notexecute_query_batch), so the connection goes back to the pool in between. Run 2 can land on a different connection and show pre-transaction data; run 3 reportsthere is no transaction in progresswhile the real transaction stays open elsewhere.It is also a hazard for unrelated queries. The pool is configured with
RecyclingMethod::Fast(src/client.rs:376), which resets nothing, so the stranded connection returns to the pool still inside the transaction, holding its locks, and the next borrower runs inside it.Fix
When the host sends a
session_id, a batch that leaves an explicit transaction open keeps its connection pinned to that session, so the next batch from the same session continues the same transaction. Committing, rolling back, an explicitrelease_session, shutdown, or 30 minutes idle releases it. A pinned connection is always rolled back before it goes back to the pool, so an open transaction can no longer leak into an unrelated query.Pinning is lazy: a session that never opens a transaction holds nothing and behaves exactly as before.
transaction_effect()inhandlers/query.rs(BEGIN,START TRANSACTION,COMMIT,END,ROLLBACK,ABORT,PREPARE TRANSACTION, and... AND CHAIN), next to the existingstrip_leading_sql_commentsit builds on. It reads leading keywords only, soBEGINin a string literal, a later clause, or a PL/pgSQLDO $$ BEGIN … END $$body is not mistaken for transaction control, andROLLBACK TO SAVEPOINTcorrectly leaves the transaction open.session.rs: the registry, theROLLBACK-before-release rule, andsweep_idle(), run from the existing 10-minuterun_pool_cleanuptimer.run_batch_in_session, extracted fromexecute_query_batch, reuses the session's pinned connection and re-pins it if a transaction is still open afterwards.execute_querygoes through it as a one-statement batch, so a single statement and a batch take the same session path. Without that, the one-at-a-time workflow above is exactly the case that stays broken. A failed ordinary statement leaves the transaction state alone (it is aborted but still open, so the user canROLLBACKfrom the same tab). A failedCOMMITis different: PostgreSQL has already rolled the transaction back, so the session is released.release_sessionRPC method;shutdownreleases every pinned session.Wire compatibility
session_idis optional on bothexecute_queryandexecute_query_batch. Without it the behaviour and the responses are byte-for-byte what they are today: a bareQueryResultand a bare array of per-statement results. Hosts older than tabularis#801 are unaffected.{ "result": {...}, "in_transaction": bool }and{ "results": [...], "in_transaction": bool }. The host side accepts both shapes permanently, and recognises the wrapper byin_transactionbeing present so a query selecting a column namedresultis not mistaken for one.release_sessionis a new method; a host that never calls it loses nothing, since the idle sweep andshutdownstill reclaim connections.One deliberate behaviour difference on a reused connection
search_pathis applied only when the connection is freshly acquired. Re-applying it to a pinned connection would runSET search_pathinside the caller's open transaction and change what the rest of that transaction sees. The consequence is that switching schema mid-transaction does not take effect until the transaction ends, which is the safer of the two behaviours, but worth a maintainer's eye.Testing
cargo test --lib: 340 pass, including 9 cases fortransaction_effectandin_transaction_after(opens / closes /AND CHAIN/ two-phase / savepoint / leading comments / data-not-keyword / PL/pgSQL body / empty input).cargo clippy --all-targets -- -D warningsandcargo fmt --checkclean.tests/live_db.rsnow hascommit_that_fails_on_a_deferred_constraint_releases_the_session(from review), but I have not run it locally.