Skip to content

Apply strip_document to BeautifulSoup Tag roots - #285

Open
ryanduguid wants to merge 1 commit into
matthewwithanm:developfrom
ryanduguid:fix-tag-root-strip-document
Open

ryanduguid wants to merge 1 commit into
matthewwithanm:developfrom
ryanduguid:fix-tag-root-strip-document

Conversation

@ryanduguid

Copy link
Copy Markdown

Summary

Apply strip_document to completed BeautifulSoup Tag conversion, addressing the boundary-newline asymmetry reported in #209. Reuse the existing newline-only policy after normal root processing when cached [document] conversion is available.

process the supplied root normally
if it is a non-document Tag and document conversion is available:
    apply the built-in strip_document policy to the result

The added finalisation does not invoke the resolved document callback. Normal tag dispatch and document callback ownership remain in place; the README describes the proposed Tag contract.

Evidence

  • Before: Of 73 new parametrised cases, 43 fail and 30 pass against the original runtime, with no collection errors. The reported Tag root retains boundary newlines.
    After: All 156 tests pass on Linux Python 3.8.20 and Windows Python 3.14.8. Coverage includes exclusions, custom callbacks, cached None, method aliases, significant whitespace and original tree relationships.
  • Configured Linux tox, python -m build -nwsx ., mypy . and mypy --strict tests/types.py pass. Windows pytest, configured flake8 and README lint pass.
  • All 73 new cases pass at BeautifulSoup 4.9.0/six 1.15.0 and against the built wheel installed outside the source checkout. The build retains the existing version-of-None warning.

Merge danger

Door: two-way. Blast radius: fragment consumers.

This proposes a compatibility change: Tag boundaries, including custom converter output, now follow strip_document; an invalid mode raises after Tag processing when document conversion is available. The extra cached document lookup, including None, can change later document behaviour with stateful resolvers. strip_document=None preserves boundary text but still performs that lookup. Maintainer acceptance of these semantics remains necessary.

Unverified

The full dependency-floor suite cannot collect the existing test_escaping.py warning import on both the base and initial candidate. Other Python/Windows versions, macOS, alternative parsers, Windows packaging/types, independent execution of the verification producers, additional exception-precedence/custom-unavailable-invalid-mode probes, scanner parser coverage, hosted CI and downstream reliance remain unverified. No performance improvement is claimed.

Reuse the document newline policy after normal Tag processing when
document conversion is available, and document the cache and validation
effects. Cover exclusions, callbacks, resolver history and tree identity.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant