Skip to content

Fix panic on modules declaring their encoding as e.g. latin-1 - #322

Open
pylaterreur wants to merge 2 commits into
python-grimp:mainfrom
pylaterreur:fix-encoding-declarations
Open

pylaterreur wants to merge 2 commits into
python-grimp:mainfrom
pylaterreur:fix-encoding-declarations

Conversation

@pylaterreur

@pylaterreur pylaterreur commented Oct 4, 2026 •

Copy link
Copy Markdown

Fixes #321.

Since 3.10, build_graph panics with TypeError('function takes exactly 5 arguments (1 given)') on:

  • modules that declare their encoding with a name Python accepts but that isn't a WHATWG label, such as latin-1, latin_1, utf_8, utf-8-sig, euc_jp or iso8859_15;
  • modules that can't be decoded at all.

Changes:

  • RealBasicFileSystem::read normalizes the declared name like Python does (_get_normal_name in Lib/tokenize.py, and _ vs -) when encoding_rs doesn't recognize it as is. Names that encoding_rs already knows are looked up as before.
  • Decoding failures are raised as UnicodeError, with the existing messages that name the file, e.g. Failed to decode file /path/to/pkg/broken.py as UTF-8: invalid utf-8 sequence of 1 bytes from index 5. Before, they were created as UnicodeDecodeError, which can't be created from a message alone.
  • FileSystem::read returns Rust errors (new GrimpError::FileNotFound and GrimpError::UndecodableFile variants), which are converted to FileNotFoundError and UnicodeError where they reach Python, like the other GrimpError variants. scan_for_imports propagates them instead of unwrapping.

The new functional tests use tmp_path, as each case needs its own package: one test per encoding name above, and one per kind of decoding failure (invalid UTF-8, invalid for the declared encoding, unknown encoding).

This isn't on a hot path: names that encoding_rs knows take the same route as before, and only the error handling changed otherwise.

Encoding names that Python accepts but that still aren't WHATWG labels after normalizing, such as cp932 or mac_roman, now raise a clear UnicodeError rather than panicking. Supporting them would need a fallback to Python's codecs.

  • Add tests for the change. In general, aim for full test coverage at the Python level. Rust tests are optional.
  • Add any appropriate documentation. (None needed.)
  • Add a summary of changes to the latest section at the top of CHANGELOG.rst. (If it's not there, add it.)
  • Add your name to AUTHORS.rst.
  • Run just full-check. (I ran the lint (Python, plus Rust with the pinned 1.97.0 toolchain), the docs build, cargo test, and the Python tests on 3.10 to 3.14. I didn't run them on 3.14t, 3.15 or 3.15t.)

🤖 Generated with Claude Code

https://claude.ai/code/session_01GwUvhmzgfkiZ6vo4uGyesN

Modules can declare their encoding (PEP 263) using any name that Python
accepts, such as `latin-1`, `utf_8` or `euc_jp`. But encoding_rs only knows
the WHATWG labels (`latin1`, `utf-8`, `euc-jp`...), so these names weren't
recognized, and building the graph panicked with:

    pyo3_runtime.PanicException: called `Result::unwrap()` on an `Err`
    value: PyErr { type: <class 'TypeError'>, value: TypeError('function
    takes exactly 5 arguments (1 given)'), traceback: None }

Normalize the declared name like Python does before looking it up.

Modules that really can't be decoded also caused this panic, as the error
was created as a UnicodeDecodeError, which can't be created from just a
message, and was then unwrapped. Raise it as a UnicodeError instead, which
keeps the message naming the file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GwUvhmzgfkiZ6vo4uGyesN
@codspeed

codspeed Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 26 untouched benchmarks
⏩ 23 skipped benchmarks1


Comparing pylaterreur:fix-encoding-declarations (06db270) with main (676e490)

Open in CodSpeed

Footnotes

  1. 23 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

The previous commit stored file read errors in GrimpError as a PyErr.
GrimpError's Debug and Display implementations then call into Python, so
the Rust test binaries started linking libpython, and `just test-rust`
failed wherever libpython isn't on the loader path:

    error while loading shared libraries: libpython3.13.so.1.0: cannot
    open shared object file: No such file or directory

Make FileSystem::read return Rust errors instead (GrimpError::FileNotFound
and GrimpError::UndecodableFile), and convert them to FileNotFoundError
and UnicodeError where they reach Python, like the other GrimpError
variants. The exception types and messages are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GwUvhmzgfkiZ6vo4uGyesN
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

build_graph panics on modules declaring their encoding as latin-1, utf_8, etc.

1 participant