Fix the per-query memory leak in the C accelerator - #136
Merged
Merged
Conversation
kesmit13
requested review from
mgiannakopoulos,
pmishchenko-ua and
volodymyr-memsql
as code owners
September 28, 2026 15:01
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved lifetime-safety, allocation-error handling, and test configuration issues remain.
Review effort: Lite
Findings: 1
Open (2)
What changed in this PR
Fixes per-query memory leaks in the C accelerator while preserving struct-sequence row lifetime.
Changes:
- Releases temporary UTF-8 and per-column allocations.
- Adds ownership handling for struct-sequence field names.
- Adds leak and row-lifetime regression tests.
| File | Description |
|---|---|
singlestoredb/tests/test_accel_leaks.py |
Adds memory-leak and lifetime regression coverage. |
accel.c |
Updates allocation cleanup and struct-sequence ownership. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
kesmit13
added a commit
that referenced
this pull request
Sep 29, 2026
Address review feedback on #136. The capsule owning a struct sequence type's field name storage was stored as __singlestoredb_fields__ on the type, where user code could delete or replace it. That decrefs the capsule and frees the names while the type and its existing rows still point at them, so a later repr(row) reads freed memory -- confirmed as a UnicodeDecodeError off garbage bytes. Ownership now lives in a module-private dict inside the extension, keyed by a weak reference to the type. The weakref callback drops the entry, and with it the capsule, when the type is collected, so there is one free path and no Python-reachable way to trigger it early. Rows are instances of a heap type and hold a reference to it, so the type still outlives every row. Also read the already-parsed pure_python option in the tests instead of parsing SINGLESTOREDB_PURE_PYTHON with int(), which raised at collection time for valid values such as `true`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_PyUnicode_AsUTF8 took a new reference to the encoded bytes and never dropped it, so every query leaked one block per column plus one for the encoding errors string -- around 50 MB per 3000 wide queries, which is enough to OOM a long-running service. Fixes #135. The per-column encodings were leaking their C copies too: State_clear_fields freed the array and not the entries in it. The struct sequence field names cannot be freed the same way. The type stores the name pointers rather than copying the strings, and reads them again when a row is repr'd, so they have to outlive every row rather than the State that built them -- freeing them here would trade the leak for a use-after-free. They now belong to a capsule in the type's own dict, which frees them when the last reference to the type goes away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Address review feedback on #136. The capsule owning a struct sequence type's field name storage was stored as __singlestoredb_fields__ on the type, where user code could delete or replace it. That decrefs the capsule and frees the names while the type and its existing rows still point at them, so a later repr(row) reads freed memory -- confirmed as a UnicodeDecodeError off garbage bytes. Ownership now lives in a module-private dict inside the extension, keyed by a weak reference to the type. The weakref callback drops the entry, and with it the capsule, when the type is collected, so there is one free path and no Python-reachable way to trigger it early. Rows are instances of a heap type and hold a reference to it, so the type still outlives every row. Also read the already-parsed pure_python option in the tests instead of parsing SINGLESTOREDB_PURE_PYTHON with int(), which raised at collection time for valid values such as `true`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
kesmit13
force-pushed
the
fix-accel-utf8-leak
branch
from
September 29, 2026 13:15
c0cccc6 to
edee40e
Compare
The namedtuples subtest of test_no_leak_per_query failed in CI while passing locally, reporting a steady 15.005 retained blocks per query. That is not our allocation. coverage.py's sys.monitoring backend keeps every code object it ever sees alive forever, deliberately, keyed by id() -- see code_objects in coverage/sysmon.py. collections.namedtuple compiles a fresh __new__ on every call and the accelerator builds one Row class per query, so a --cov run, which is what code-check.yml does, retains a code object per query with nothing wrong. pandas' DataFrame.itertuples builds its class the same way with no cache, so this is the tracer's accounting rather than a C API artifact or ours to fix. The leak in issue #135 cost one allocation per column per query, and the accelerator now measures flat at 15.005 blocks per query for both a 10-column and a 100-column result -- a constant that does not move when the width grows tenfold was never that bug. So assert the property the bug actually had: difference a narrow and a wide query, which cancels every per-query cost that is flat in the column count. The tracer overhead is flat, measured unchanged from 5 to 200 columns, so it cancels exactly rather than approximately. Keep a separate absolute per-query check, since differencing two widths cannot see something leaked once per query, and fund the namedtuples budget with an overhead figure measured at runtime so it stays tight when nothing is tracing. Verified both ways: the differential assertion still reports 1.0 blocks per column for tuples, dicts and namedtuples and 2.0 for structsequences when built against accel.c as of 4e348a8, twenty times the threshold, and the whole file passes on 3.11 and on 3.14.7 with and without --cov. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_PyUnicode_AsUTF8 returns a calloc'd copy now, so it can fail. Three of its four call sites check for that; the per-column one in State_init did not, and NULL is the binary-column sentinel every reader of encodings[] goes by (accel.c:1944, 1974, 2159). A failed allocation for a text column would therefore have decoded that column as binary, with the PyErr_NoMemory only surfacing when the function exited. Split the ternary into an explicit branch so the failure is distinguishable from the sentinel and reaches the error label. The !py_encoding arm it replaces was dead -- the PyTuple_GetItem above already bails on NULL. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Fixes #135.
_PyUnicode_AsUTF8inaccel.ctook a new reference to the encodedbytesand never dropped it, so every query leaked one block per columnplus one for the encoding errors string. Measured on a 200-column query,
3000 iterations: ~50 MB of RSS growth, and 201 blocks per query —
matching the numbers in the issue.
There were four call sites, not the three the issue lists;
get_numpy_col_typeleaks the samebytes(it does free its C copy).What changed
Py_DECREFs thebyteson every path, and checkscallocbefore writing through it.Py_LIMITED_APIstays at 3.8, sothe hand-rolled copy stays as well —
PyUnicode_AsUTF8AndSizeis notavailable to us until the floor moves to 3.10.
State_clear_fieldsfreed the array and not the entries in it.PyStructSequence_NewTypestores the name pointers rather than copyingthe strings (
initialize_membersinObjects/structseq.c), andstructseq_reprreads them again later. Rows can outlive theStatethat built them, so freeing the names in
State_clear_fieldswouldhave traded the leak for a use-after-free in
repr(row). They nowbelong to a capsule in the type's own dict, freed when the last
reference to the type goes away.
State_clear_fieldsonly frees themon the paths where the capsule never took ownership, and now does so
after
Py_CLEAR(self->structsequence)rather than before.Verification
singlestoredb/tests/test_accel_leaks.pyruns a 100-column query 200times per
results_typeand asserts thesys.getallocatedblocks()delta stays under 5 blocks/query. On this branch it measures 0.00; on
mainit fails at 101.005. Skipped when the extension is unavailable,under
SINGLESTOREDB_PURE_PYTHON=1, and on HTTP connections.The second test holds struct sequence rows, churns the
State, thenreads every field and the repr — it passes on
maintoo (the leak iswhat kept those names alive), and exists to stop this fix regressing
into a use-after-free.
Before and after, 3000 iterations:
mainFull non-management suite: 851 passed, 19 skipped.
🤖 Generated with Claude Code
Note
Medium Risk
Changes native extension memory ownership and struct-sequence field lifetime; incorrect freeing could cause use-after-free in row repr, though new tests target that regression.
Overview
Fixes per-query memory leaks in the C MySQL accelerator (
accel.c), including the issue #135 pattern of ~one allocated block per result column._PyUnicode_AsUTF8now always releases the intermediatebytesobject (previously leaked on every call) and handlescallocfailure on the error path. Call sites for column encodings and struct-sequence field names propagate allocation failures instead of leaving bad state.State_clear_fieldsfrees each per-column encoding C string, not only the pointer array. Struct-sequence field name strings are no longer torn down with query state: they move into a module registry (weak ref to the row type → capsule with destructor) so names stay valid forrepron rows that outlive the fetch state.Adds
test_accel_leaks.py:sys.getallocatedblocks()budgets for per-column and per-query retention acrossresults_typemodes, plus struct-sequence lifetime checks after state churn and type-dict stripping.Reviewed by Cursor Bugbot for commit 2a1b3b6. Bugbot is set up for automated code reviews on this repo. Configure here.