Skip to content

feat(filter): skip model weight files by default - #95

Merged
matiasdaloia merged 1 commit into
mainfrom
feat/skip-model-weights
Sep 25, 2026
Merged

matiasdaloia merged 1 commit into
mainfrom
feat/skip-model-weights

Conversation

@matiasdaloia

Copy link
Copy Markdown
Contributor

Why

wfp fingerprinting reads each collected file whole into memory. The default filters skipped no model-weight format and impose no size cap, so a repository holding a few weights (for example a 256 MB .safetensors, a 239 MB pytorch_model.bin and an 89 MB .gguf) pushed the fingerprinting process past 2 GB of resident memory. Snippet fingerprints of weight bytes never produce a useful match, so the cost buys nothing.

What

  • defaultSkippedExts gains a commented group of model-weight extensions: .safetensors .gguf .ggml .bin .onnx .pt .pth .ckpt .h5 .hdf5 .keras .tflite .pb .npy .npz .pkl .joblib .mlmodel .msgpack .ot .caffemodel .nemo. None was already in the list. Matching is case-insensitive through the existing extension set.
  • .bin is included deliberately, so any other .bin file is skipped too.
  • These are built-in file rules, so they apply to Scanning, Fingerprinting and Dependencies. No dependency manifest in pkg/manifests uses one of these extensions, so KeepManifests behaviour is unchanged.
  • Callers that need these files must collect them separately, or turn off the built-in file rules (--all-extensions, BuiltinFileRules: false).
  • CHANGELOG (Unreleased / Changed) and the "Skipping files" section of CLIENT_HELP.md are updated.

Verification

  • New TestModelWeightsAreSkipped runs the real Collect walk with Scanning(nil) over every extension in lower and upper case beside a train.py, and asserts only train.py is collected and the skip count is 44.
  • Break-check: removing .nemo from the list turns the test red (collected [UPPER.NEMO lower.nemo train.py], want [train.py]); restoring it turns it green.
  • make test passes; make lint reports 0 issues.

Fingerprinting reads each file whole, so a tree holding a few model
weights of hundreds of megabytes each drove the process to gigabytes of
memory, and snippet fingerprints of weight bytes never produce a useful
match.

The default file rules shared by the Scanning, Fingerprinting and
Dependencies profiles now skip .safetensors, .gguf, .ggml, .bin, .onnx,
.pt, .pth, .ckpt, .h5, .hdf5, .keras, .tflite, .pb, .npy, .npz, .pkl,
.joblib, .mlmodel, .msgpack, .ot, .caffemodel and .nemo. Callers that
need these files must collect them separately or disable the built-in
file rules.
@matiasdaloia
matiasdaloia merged commit 0028de1 into main Sep 25, 2026
5 checks passed
@matiasdaloia
matiasdaloia deleted the feat/skip-model-weights branch September 25, 2026 22:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant