Repository navigation
Fix UnicodeDecodeError reading non-UTF-8 .py module files - #281
Conversation
Fall back to cp1252 (Windows ANSI) when a .py file is not valid UTF-8 so that metadata extraction does not crash circup. Fixes adafruit#222.
|
Would |
|
I tried the |
The dunder regex only reads ASCII, so replacing undecodable bytes covers cp1252, latin-1 and mac-roman alike and drops the second open() and the extra pylint disable. The existing extract_metadata test asserts the open() call signature, so it moves to the new one.
|
Done in a498363. One thing your suggestion could not see: On AI usage: yes, this was AI-assisted. I draft and run the tests with Claude, and review and verify everything before it goes up. Happy to record that in the PR description if you would rather have it there. |
mikeysklar
left a comment
There was a problem hiding this comment.
Thanks! No need to change this one, but please note AI usage in the PR description going forward.
extract_metadataincircup/shared.pyopened.pyfiles as UTF-8 and let the decode errorpropagate, so a single file saved in a Windows ANSI encoding, one with curly quotes in a comment for
instance, crashed the whole scan with
UnicodeDecodeErrorrather than being skipped or read.This opens those files with
errors="replace". The dunder regex only reads the ASCII parts, soreplacing undecodable bytes is enough for cp1252, latin-1 and mac-roman alike, and a file that
previously crashed now yields its
__version__and__repo__.Files that already decoded as UTF-8 are unaffected.
The new test writes a cp1252 encoded file with a curly quote and asserts the metadata comes back
rather than an exception. The existing
test_extract_metadata_pythonasserts theopen()callsignature, so it moves to the new one.
Fixes #222