Skip to content

PDF: batch table scripts and geometry evidence - #28

Merged
myhloli merged 3 commits into
codex/rust-pdf-stage20-geometry-materializationfrom
codex/rust-pdf-stage21-table-geometry-batching
Sep 29, 2026
Merged

myhloli merged 3 commits into
codex/rust-pdf-stage20-geometry-materializationfrom
codex/rust-pdf-stage21-table-geometry-batching

Conversation

@myhloli

@myhloli myhloli commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Summary

  • Stack on Stage 20 d504c5e; private protocol advances 27 → 28.
  • Reuse page-owned NativeScriptEvidence for table-cell script classification with bounded batches and conservative whole-cell reference fallback.
  • Cache marker-line validation once per rule-candidate context instead of once per corridor.
  • Add page-level NativeGeometryEvidence and a document-wide native run table so full-document geometry-risk admission consumes line member indices rather than repacking every character through Python dictionaries.
  • Preserve public auto|python|rust compute and auto|legacy|session render selection, public SDK shape, output schema, and MinerU behavior.

Correctness

  • Public replay: 32 PDFs / 299 pages, two candidate runs per document; ModelJson, MiddleJson, assets, and diagnostics matched Stage 20 exactly.
  • MinerU dev@f504cff: 31 eligible Flash text outputs and 32 actual medium shared outputs all matched Stage 20.
  • Rust/session full tests: 5465 passed, 14 skipped.
  • Python/legacy full tests: 4756 passed, 723 skipped.
  • Cargo workspace tests, Clippy -D warnings, rustfmt, Ruff check, and changed-file formatting passed.
  • ABI3 wheel verified in a CPython 3.14 isolated dependency environment. Source and wheel extension SHA-256 both 6266f409337d742153dc5bc0727d3b28f32a8f4cb210d8a5ea6e06cc650d3deb; DocVortex core and MinerU Flash smoke outputs matched for auto and rust.

Diagnostics

  • Public correctness table script path: 368 batches, 27,602 owned lines, 932 whole-cell fallbacks.
  • Public correctness geometry evidence: 598 pages, 34,990 lines, 0 fallbacks.
  • One-pass 32-PDF profile: geometry evidence 299 pages / 17,495 lines; table script path 184 batches / 13,801 lines.

Measured performance

One warmup plus five hot runs per document, alternating baseline/candidate execution where specified, with isolated process-tree RSS sampling.

Chain Stage 20 Stage 21 First-pass reduction Max first time Max first RSS
DocVortex public parse, 32 PDFs 15.943439 s 15.513413 s 2.70% 1.02077 1.03600
MinerU Flash text, 31 PDFs 15.746155 s 15.304521 s 2.80% 1.02846 1.04413
MinerU medium shared, 32 PDFs 16.010728 s 17.845207 s -11.46% 1.58188 1.00490

Four shared samples exceeded 5% in the first pass and were reverse-order retested:

  • small_ocr.pdf: 1.09802 → 0.98395
  • engineering_process_restrictions_table.pdf: 1.51380 → 0.99771
  • pollutant_discharge_tables.pdf: 1.58188 → 0.88237
  • quarterly_report_financial_tables.pdf: 1.12287 → 0.90017

None remained above 5%. As a diagnostic aggregate replacing only those four triggered samples with reverse retests, shared time was 16.174975 s → 15.777589 s (2.46% improvement); this does not replace the first-pass total above. All compared outputs were equal.

Ten table-heavy focus documents reduced wall time from 11.889497 s to 10.338891 s (13.03%). The three selected hotspots fell from 2.979244 s to 0.117086 s, below the 2.80 s stage gate:

  • _cell_script_roles: 1.969974 s → 0.040009 s
  • _prepare_table_core_rows: 0.542125 s → 0.026794 s
  • _document_requires_full_geometry: 0.467145 s → 0.050283 s

The original 2× target remains incomplete. Remaining Stage 22 candidates include table candidate detection/materialization, lane inference, geometry-plan sample construction, and owned table candidate merging.

Not done

  • No merge, release, version bump, PyPI publication, or downstream dependency synchronization.
  • MinerU is integration-only; its untracked examples/ directory was preserved.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T16:18:12.139897Z 4c77121 New commits
🔒 Security Review ✅ Completed 2026-09-29T15:21:00.041360Z 8387767 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8387767878

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/docvortex/analyzers/native/pdf/char_geometry.py
Owned batches intern run ids in NativeGeometryRuns while per-line fallback
reallocated ids from zero in a Python dict; mixing both inside one
NativeGeometryRisk accumulator merged unrelated runs and split identical
ones, silently skewing the full-document risk verdict. Documents that opt
into owned evidence now restart the whole calculation on the reference
path as soon as any line misses, matching the documented whole-document
fallback semantics.
@myhloli
myhloli merged commit 90a0e8a into codex/rust-pdf-stage20-geometry-materialization Sep 29, 2026
21 checks passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4c771211d4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +496 to +499
let rounded_size: f64 =
round.call1((font.size * 4.0,))?.extract::<i64>()? as f64 / 4.0;
let weight: i64 = round.call1((f64::from(font.weight) / 100.0,))?.extract()?;
let value = (family, rounded_size, font.flags, weight as i32);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the unrounded font size for style-risk checks

For fonts whose size is not quarter-point aligned, this stores only rounded_size, and NativeGeometryEvidence.add_line() later passes that same value both as the run-key component and as the font-size sample. The Python reference intentionally rounds only the run key while returning the raw size from _font_run_key() for _style_line_is_inflated; near the 4-point eligibility or 1.5× inflation thresholds, documents with repeated lines across pages can therefore get a different style-risk decision and unnecessary or missing geometry repairs. Store the raw size separately and use the rounded value only when interning the run key.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant