Skip to content

Build pydantic validators on first use to reduce memory - #159

Open
dirkjanfaber wants to merge 1 commit into
flexiblepower:mainfrom
dirkjanfaber:defer-model-building
Open

dirkjanfaber wants to merge 1 commit into
flexiblepower:mainfrom
dirkjanfaber:defer-model-building

Conversation

@dirkjanfaber

Copy link
Copy Markdown

Reduce memory use by building pydantic validators on first use

Importing s2python currently builds the pydantic validators and serializers of all 126 models (65 generated + 61 message classes). That costs about 10 MB of memory per process, while applications typically use only a handful of messages.
On embedded devices (like our GX devices) running several services that use s2python, this adds up.

Changes

  • The generated models get an S2BaseModel base class with defer_build=True (via --base-class in ci/generate_s2.sh), so a model is built the first time it is used.
  • Fixes a latent bug that made this impossible: the message classes took model_config and model_fields[...] from the generated models by reference. model_config["validate_assignment"] = True changed the config of the generated
    model, and pydantic rewrote the reused FieldInfo objects (e.g. the generated DDBCOperationMode.Id ended up typed uuid.UUID instead of ID). This went unnoticed while the generated models were built before being modified, but with deferred building their validation changed. The new copy_config() and copy_field() helpers copy them instead.

Verification

  • The JSON schemas of all 63 public models are identical before and after.
  • After importing, none of the generated models is modified (before: 61).
  • Unit tests pass with pydantic 2.8.2 and 2.13.5; mypy, pyright and pylint are clean (mypy even reports one error fewer than before).
  • Memory with the handshake and OMBC messages in use, on an ARM device (Python 3.12, pydantic 2.7.4): 10.3 MB -> 6.7 MB.

Notes

  • gen_s2.py contains a manual change that is not in the specification: populate_by_name=True on Transition (Fix for transition argument #113). Regenerating would drop it; this PR keeps it. It should probably move into the specification.
  • The diff of gen_s2.py only changes the base class. A full regeneration with datamodel-code-generator 0.32.0 produces the same code apart from formatting.
  • Kept the inconsistency of LF and CRLF to keep the diffs minimal (the 13 files in ddbc/ and ombc/ use CRLF line endings, while the rest of the repo uses LF, and there's no .gitattributes.)

Importing s2python built the validators and serializers of all 126
models, which costs about 10 MB of memory, while applications typically
use only a few of them.

- Generate the models with an S2BaseModel base class that sets
  defer_build=True.
- The message classes copied config and fields from the generated
  models by reference, so pydantic modified the generated models when
  the subclasses were created. That went unnoticed as long as the
  generated models were built first, but broke with deferred building.
  Copy them with the new copy_config() and copy_field() helpers instead.

The JSON schemas of all public models are unchanged. On an ARM device
with pydantic 2.7.4 memory use drops from 10.3 MB to 6.7 MB.
@jorritn

jorritn commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

The rationale totally makes sense to me and I like the defer_build approach to achieve it. I would like to have @sebastiaan-la-fleur have a look at it because he set up the whole pydantic based class generations. What do you think of decreasing memory consumption this way and do you see any unwanted side effects?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants