Repository navigation
PDF: batch table scripts and geometry evidence - #28
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8387767878
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Owned batches intern run ids in NativeGeometryRuns while per-line fallback reallocated ids from zero in a Python dict; mixing both inside one NativeGeometryRisk accumulator merged unrelated runs and split identical ones, silently skewing the full-document risk verdict. Documents that opt into owned evidence now restart the whole calculation on the reference path as soon as any line misses, matching the documented whole-document fallback semantics.
90a0e8a
into
codex/rust-pdf-stage20-geometry-materialization
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 4c771211d4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| let rounded_size: f64 = | ||
| round.call1((font.size * 4.0,))?.extract::<i64>()? as f64 / 4.0; | ||
| let weight: i64 = round.call1((f64::from(font.weight) / 100.0,))?.extract()?; | ||
| let value = (family, rounded_size, font.flags, weight as i32); |
There was a problem hiding this comment.
Preserve the unrounded font size for style-risk checks
For fonts whose size is not quarter-point aligned, this stores only rounded_size, and NativeGeometryEvidence.add_line() later passes that same value both as the run-key component and as the font-size sample. The Python reference intentionally rounds only the run key while returning the raw size from _font_run_key() for _style_line_is_inflated; near the 4-point eligibility or 1.5× inflation thresholds, documents with repeated lines across pages can therefore get a different style-risk decision and unnecessary or missing geometry repairs. Store the raw size separately and use the rounded value only when interning the run key.
Useful? React with 👍 / 👎.
Summary
d504c5e; private protocol advances 27 → 28.NativeScriptEvidencefor table-cell script classification with bounded batches and conservative whole-cell reference fallback.NativeGeometryEvidenceand a document-wide native run table so full-document geometry-risk admission consumes line member indices rather than repacking every character through Python dictionaries.auto|python|rustcompute andauto|legacy|sessionrender selection, public SDK shape, output schema, and MinerU behavior.Correctness
dev@f504cff: 31 eligible Flash text outputs and 32 actual medium shared outputs all matched Stage 20.-D warnings, rustfmt, Ruff check, and changed-file formatting passed.6266f409337d742153dc5bc0727d3b28f32a8f4cb210d8a5ea6e06cc650d3deb; DocVortex core and MinerU Flash smoke outputs matched forautoandrust.Diagnostics
Measured performance
One warmup plus five hot runs per document, alternating baseline/candidate execution where specified, with isolated process-tree RSS sampling.
Four shared samples exceeded 5% in the first pass and were reverse-order retested:
small_ocr.pdf: 1.09802 → 0.98395engineering_process_restrictions_table.pdf: 1.51380 → 0.99771pollutant_discharge_tables.pdf: 1.58188 → 0.88237quarterly_report_financial_tables.pdf: 1.12287 → 0.90017None remained above 5%. As a diagnostic aggregate replacing only those four triggered samples with reverse retests, shared time was 16.174975 s → 15.777589 s (2.46% improvement); this does not replace the first-pass total above. All compared outputs were equal.
Ten table-heavy focus documents reduced wall time from 11.889497 s to 10.338891 s (13.03%). The three selected hotspots fell from 2.979244 s to 0.117086 s, below the 2.80 s stage gate:
_cell_script_roles: 1.969974 s → 0.040009 s_prepare_table_core_rows: 0.542125 s → 0.026794 s_document_requires_full_geometry: 0.467145 s → 0.050283 sThe original 2× target remains incomplete. Remaining Stage 22 candidates include table candidate detection/materialization, lane inference, geometry-plan sample construction, and owned table candidate merging.
Not done
examples/directory was preserved.