feat(filter): skip model weight files by default - #95
Merged
Merged
Conversation
Fingerprinting reads each file whole, so a tree holding a few model weights of hundreds of megabytes each drove the process to gigabytes of memory, and snippet fingerprints of weight bytes never produce a useful match. The default file rules shared by the Scanning, Fingerprinting and Dependencies profiles now skip .safetensors, .gguf, .ggml, .bin, .onnx, .pt, .pth, .ckpt, .h5, .hdf5, .keras, .tflite, .pb, .npy, .npz, .pkl, .joblib, .mlmodel, .msgpack, .ot, .caffemodel and .nemo. Callers that need these files must collect them separately or disable the built-in file rules.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
wfpfingerprinting reads each collected file whole into memory. The default filters skipped no model-weight format and impose no size cap, so a repository holding a few weights (for example a 256 MB.safetensors, a 239 MBpytorch_model.binand an 89 MB.gguf) pushed the fingerprinting process past 2 GB of resident memory. Snippet fingerprints of weight bytes never produce a useful match, so the cost buys nothing.What
defaultSkippedExtsgains a commented group of model-weight extensions:.safetensors .gguf .ggml .bin .onnx .pt .pth .ckpt .h5 .hdf5 .keras .tflite .pb .npy .npz .pkl .joblib .mlmodel .msgpack .ot .caffemodel .nemo. None was already in the list. Matching is case-insensitive through the existing extension set..binis included deliberately, so any other.binfile is skipped too.Scanning,FingerprintingandDependencies. No dependency manifest inpkg/manifestsuses one of these extensions, soKeepManifestsbehaviour is unchanged.--all-extensions,BuiltinFileRules: false).Unreleased / Changed) and the "Skipping files" section of CLIENT_HELP.md are updated.Verification
TestModelWeightsAreSkippedruns the realCollectwalk withScanning(nil)over every extension in lower and upper case beside atrain.py, and asserts onlytrain.pyis collected and the skip count is 44..nemofrom the list turns the test red (collected [UPPER.NEMO lower.nemo train.py], want [train.py]); restoring it turns it green.make testpasses;make lintreports 0 issues.