Build pydantic validators on first use to reduce memory - #159
Open
dirkjanfaber wants to merge 1 commit into
Open
dirkjanfaber wants to merge 1 commit into
dirkjanfaber wants to merge 1 commit into
Conversation
Importing s2python built the validators and serializers of all 126 models, which costs about 10 MB of memory, while applications typically use only a few of them. - Generate the models with an S2BaseModel base class that sets defer_build=True. - The message classes copied config and fields from the generated models by reference, so pydantic modified the generated models when the subclasses were created. That went unnoticed as long as the generated models were built first, but broke with deferred building. Copy them with the new copy_config() and copy_field() helpers instead. The JSON schemas of all public models are unchanged. On an ARM device with pydantic 2.7.4 memory use drops from 10.3 MB to 6.7 MB.
Contributor
|
The rationale totally makes sense to me and I like the defer_build approach to achieve it. I would like to have @sebastiaan-la-fleur have a look at it because he set up the whole pydantic based class generations. What do you think of decreasing memory consumption this way and do you see any unwanted side effects? |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reduce memory use by building pydantic validators on first use
Importing s2python currently builds the pydantic validators and serializers of all 126 models (65 generated + 61 message classes). That costs about 10 MB of memory per process, while applications typically use only a handful of messages.
On embedded devices (like our GX devices) running several services that use s2python, this adds up.
Changes
S2BaseModelbase class withdefer_build=True(via--base-classinci/generate_s2.sh), so a model is built the first time it is used.model_configandmodel_fields[...]from the generated models by reference.model_config["validate_assignment"] = Truechanged the config of the generatedmodel, and pydantic rewrote the reused
FieldInfoobjects (e.g. the generatedDDBCOperationMode.Idended up typeduuid.UUIDinstead ofID). This went unnoticed while the generated models were built before being modified, but with deferred building their validation changed. The newcopy_config()andcopy_field()helpers copy them instead.Verification
Notes
gen_s2.pycontains a manual change that is not in the specification:populate_by_name=TrueonTransition(Fix for transition argument #113). Regenerating would drop it; this PR keeps it. It should probably move into the specification.gen_s2.pyonly changes the base class. A full regeneration with datamodel-code-generator 0.32.0 produces the same code apart from formatting.ddbc/andombc/use CRLF line endings, while the rest of the repo uses LF, and there's no.gitattributes.)