Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 46 additions & 0 deletions docs/rust-pdf-stage22-evidence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# PDF 原生内核阶段 22:表格网格复用与栏带扫描降重

## 实现边界

本阶段以 `main@878870f`(包含 Stage 21 合并与开放 Path 修复)为基线,只减少既有 Python 边界的重复工作;私有协议保持 28,不修改 Rust ABI、公开 SDK、`auto|python|rust` 计算选择、`auto|legacy|session` 渲染选择或输出 schema。

表格检测侧新增同一候选上下文的网格连通分量冻结:延迟候选先计算一次 `_connected_rule_grid_components()`,闭合网格检测和 owned merge 只复用该分量重建 bbox。直接调用旧内部函数时仍按原签名即时计算;owned merge 对复用分量执行普通 `_LocalAxisLine` 与有限 bbox 校验,特殊输入仍回退完整参考路径。marker-safe 上下文校验改为单次顺序扫描,同一字符/字典字段不再被多层生成式重复读取,特殊 char/Bbox 仍整体回退。

栏带推断侧保留算法语义,但删除三类重复扫描:只有一个 lane 时短尾不可能跨栏,直接跳过校验和前序扫描;lane 分配前一次排序栏带,`_fits_only_one_lane_ordered()` 不再为每个 bbox 重复排序;最佳/次高覆盖在一次遍历中求得,平局仍保留首个最佳 lane。表格 cell glyph 已由 `_cell_glyphs()` 全局排序,visual row 分组后不再重复稳定排序。自有 `Bbox` 增加 exact-type 快路径,仅当内部是四个有限 float 且坐标正向时直接返回 tuple,反向、非有限、int 或子类仍走原通用转换。

## 正确性

- 32 PDF/299 页公开回放每份双跑;ModelJson、MiddleJson、素材哈希和诊断与 Stage 21 正式候选输出完全一致。
- MinerU 当前 `dev@f504cff`:31 份 eligible Flash 文本完整输出一致;32 份实际 medium shared 完整输出一致。MinerU 工作区仅保留原有未跟踪 `examples/`,未修改其源码。
- Rust/session 完整测试:5472 passed / 14 skipped。
- Python/legacy 完整测试:4762 passed / 724 skipped。
- Cargo workspace tests、Clippy `-D warnings`、rustfmt、Ruff check 均通过;本仓既有 20 个无关测试/工具文件不符合当前 ruff format,未纳入本阶段 patch,改动文件 format check 通过。
- ABI3 wheel 为 `docvortex-0.5.7-cp310-abi3-macosx_11_0_arm64.whl`。CPython 3.14.4 独立环境中 DocVortex 核心路径和 MinerU Flash 文本路径的 `auto|rust` 输出一致,协议 28。源码扩展与 wheel 扩展 SHA-256 均为 `bc42f1546006c340cb7175b5bbfa6698641a59057758c046a8f5f4c0bf5a222e`。

## 性能与资源

正式计时均为每文档一次预热、五次热运行,基线/候选执行方向逐文档交替;RSS 使用独立进程树采样。首轮任一样本耗时超过 5% 时均按相反方向复测,复测均无持续退化。结果仅代表本机语料与当前 PDFium 环境。

| 链路 | main 基线 | Stage 22 候选 | 首轮总降幅 | 首轮最大耗时比 | 首轮最大 RSS 比 |
| --- | ---: | ---: | ---: | ---: | ---: |
| DocVortex 公开 parse,32 PDF | 16.535675 s | 16.049855 s | 2.94% | 1.19798 | 1.00948 |
| MinerU Flash 文本,31 PDF | 16.379259 s | 15.990359 s | 2.37% | 1.10413 | 1.05150 |
| MinerU medium shared,32 PDF | 19.056479 s | 18.956444 s | 0.52% | 1.26994 | 1.04256 |

公开 parse 中 `demo4.pdf`、`mixed_text_layout_sample.pdf.xor` 和 `quarterly_report_financial_tables.pdf` 首轮耗时触发复测,反序比例分别为 `0.94217`、`1.04183` 和 `0.98873`,无持续退化。Flash 中 `demo4.pdf`、`annual_report_research_projects_table.pdf` 和 `fund_asset_and_transaction_tables.pdf` 触发复测,反序比例分别为 `0.96987`、`0.93490` 和 `1.01703`,无持续退化。shared 中 `mixed_elements_pages_03_06.pdf`、`small_ocr.pdf` 和 `pollutant_discharge_tables.pdf` 触发复测,反序耗时比例分别为 `1.00299`、`0.98000` 和 `0.95873`,无持续退化。

一轮 32 PDF cProfile 中,本轮针对的累计热点下降如下:

- `_detect_table_candidates()`: `6.821288 s → 4.258895 s`
- `_materialize_table_blocks()`: `5.383532 s → 3.768801 s`
- `_infer_text_lanes()`: `3.251807 s → 2.474491 s`
- `_build_rule_table_candidates()`: `3.227190 s → 2.065146 s`
- `_merge_owned_table_candidates()`: `2.223635 s → 1.366970 s`
- `_reattach_cross_lane_short_tails()`: `1.293215 s → 0.971536 s`
- `_connected_rule_grid_components()`: `0.851147 s → 0.487159 s`
- `_cell_visual_lines()`: `0.786876 s → 0.455115 s`
- `_prepare_marker_line_context()`: `0.615293 s → 0.249741 s`

原始 2 倍性能目标仍未完成。Stage 23 需重新从最终 cProfile 选择,主要候选包括 `_detect_table_candidates()` 约 `4.259 s`、`_materialize_table_blocks()` 约 `3.769 s`、`build_document_geometry_plan()` 约 `3.435 s`、`detect_pdf_text_script_lines()` 约 `2.385 s`、`_infer_text_lanes()` 约 `2.474 s`,以及 geometry canonical 样本的 `metadata()`/`_plain_source_records()` 剩余打包成本。不应把正确性通过或本阶段局部热点下降表述为 2 倍目标完成。

证据目录为 `output/pdf/native-kernel-20260930-stage22/`。本阶段不自动合并、发版或发布 PyPI。
16 changes: 15 additions & 1 deletion src/docvortex/analyzers/native/pdf/_native_table_merge.py
Original file line number Diff line number Diff line change
Expand Up @@ -164,10 +164,24 @@ def merge_owned(candidates):
return None
grids = {}
for key, (context, rows) in contexts.items():
grid_components = context.grid_components
if grid_components is not None and not (
type(grid_components) is list
and all(
type(component) is list
and all(type(rule) is rules._LocalAxisLine and _plain_box(rule.bbox) for rule in component)
for component in grid_components
)
):
grid_components = None
if context.grids is None:
context.grids = [
box
for box in rules._connected_rule_grid_bboxes(context.axis_lines, context.median_height)
for box in rules._connected_rule_grid_bboxes(
context.axis_lines,
context.median_height,
components=grid_components,
)
if not any(rules._bbox_overlap_in_smaller(box, excluded) >= 0.5 for excluded in context.excluded_bboxes)
]
if not all(_plain_box(box) for box in context.grids):
Expand Down
13 changes: 13 additions & 0 deletions src/docvortex/analyzers/native/pdf/geometry.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@


from ....schema import BBox
from ....document.pdf.text._contracts import Bbox as CharacterBbox

from .models import _AxisLine, _LocalAxisLine

Expand Down Expand Up @@ -139,6 +140,18 @@ def _rotate_origin_to_upright(
def _coerce_bbox(value: Any) -> BBox | None:
"""将任意四元 bbox 规范成非退化浮点坐标。"""

# 自有 Bbox 是热路径中的固定 slot 容器;先按原校验规则读取内部
# 四个 float,可避免通用迭代转换和重复下标调用,异常形状仍走旧路径。
if type(value) is CharacterBbox:
raw = value.bbox
if (
type(raw) is list
and len(raw) == 4
and raw[0] < raw[2]
and raw[1] < raw[3]
and all(type(item) is float and math.isfinite(item) for item in raw)
):
return (raw[0], raw[1], raw[2], raw[3])
try:
x0, y0, x1, y1 = [float(item) for item in value]
except (TypeError, ValueError):
Expand Down
91 changes: 56 additions & 35 deletions src/docvortex/analyzers/native/pdf/line_layout.py
Original file line number Diff line number Diff line change
Expand Up @@ -310,6 +310,7 @@ def _infer_text_lanes(
)
if nested_column_band is not None:
nested_lanes, band_top, band_bottom = nested_column_band
nested_ordered_lanes = sorted(nested_lanes, key=lambda lane: lane.left)
fallback_lane = _TextLane(
left=filtered_intervals[0][0],
right=filtered_intervals[0][1],
Expand All @@ -324,26 +325,17 @@ def _infer_text_lanes(
):
fallback_lane.lines.append(item)
continue
line_width = max(0.1, bbox[2] - bbox[0])
scored_lanes = [
(
max(0.0, min(bbox[2], lane.right) - max(bbox[0], lane.left)) / line_width,
lane,
)
for lane in nested_lanes
]
coverage_scores = sorted(
(coverage for coverage, _lane in scored_lanes),
reverse=True,
best_lane, best_coverage, second_coverage = _best_lane_coverage(
bbox,
nested_lanes,
)
best_coverage, best_lane = max(scored_lanes, key=lambda value: value[0])
fits_only_one_lane = _fits_only_one_lane(
fits_only_one_lane = _fits_only_one_lane_ordered(
bbox,
best_lane,
nested_lanes,
nested_ordered_lanes,
anchor_tolerance,
)
if len(coverage_scores) > 1 and coverage_scores[1] >= 0.2 and not fits_only_one_lane:
if second_coverage >= 0.2 and not fits_only_one_lane:
span_lines.append(item)
continue
if best_coverage >= 0.5 or fits_only_one_lane:
Expand Down Expand Up @@ -372,31 +364,20 @@ def _infer_text_lanes(
return lanes

lanes = [_TextLane(left=left, right=right) for left, right, _support in filtered_intervals]
ordered_lanes = sorted(lanes, key=lambda lane: lane.left)
span_lines: list[tuple[_LineItem, BBox]] = []
for item in line_geometry:
bbox = item[1]
line_width = max(0.1, bbox[2] - bbox[0])
scored_lanes = [
(
max(0.0, min(bbox[2], lane.right) - max(bbox[0], lane.left)) / line_width,
lane,
)
for lane in lanes
]
coverage_scores = sorted(
(coverage for coverage, _lane in scored_lanes),
reverse=True,
)
best_coverage, best_lane = max(scored_lanes, key=lambda value: value[0])
fits_only_one_lane = _fits_only_one_lane(
best_lane, best_coverage, second_coverage = _best_lane_coverage(bbox, lanes)
fits_only_one_lane = _fits_only_one_lane_ordered(
bbox,
best_lane,
lanes,
ordered_lanes,
anchor_tolerance,
)
# 同时覆盖两个稳定正文栏的行仍属于跨栏内容;只进入单侧栏且未越过栏沟的
# 宽正文行则回到该栏,避免窄图注把正文错误挤入 span lane。
if len(coverage_scores) > 1 and coverage_scores[1] >= 0.2 and not fits_only_one_lane:
if second_coverage >= 0.2 and not fits_only_one_lane:
span_lines.append(item)
continue
if len(lanes) == 1 or fits_only_one_lane:
Expand All @@ -423,15 +404,35 @@ def _infer_text_lanes(
return lanes


def _fits_only_one_lane(
def _best_lane_coverage(
bbox: BBox,
best_lane: _TextLane,
lanes: list[_TextLane],
) -> tuple[_TextLane, float, float]:
"""一次遍历求最佳栏覆盖与次高覆盖,平局保留原有首个最佳栏。"""

line_width = max(0.1, bbox[2] - bbox[0])
best_lane = lanes[0]
best_coverage = -1.0
second_coverage = -1.0
for lane in lanes:
coverage = max(0.0, min(bbox[2], lane.right) - max(bbox[0], lane.left)) / line_width
if coverage > best_coverage:
second_coverage = best_coverage
best_coverage = coverage
best_lane = lane
elif coverage > second_coverage:
second_coverage = coverage
return best_lane, best_coverage, second_coverage


def _fits_only_one_lane_ordered(
bbox: BBox,
best_lane: _TextLane,
ordered: list[_TextLane],
tolerance: float,
) -> bool:
"""判断宽行是否仍完整停留在某一栏及其栏沟边界以内。"""
"""在调用方预排序的栏序列上判断宽行是否只停留在一个栏内。"""

ordered = sorted(lanes, key=lambda lane: lane.left)
lane_index = ordered.index(best_lane)
if lane_index > 0 and bbox[0] < ordered[lane_index - 1].right - tolerance:
return False
Expand All @@ -440,6 +441,22 @@ def _fits_only_one_lane(
return best_lane.left - tolerance <= _bbox_center_x(bbox) <= best_lane.right + max(tolerance, bbox[2] - best_lane.right)


def _fits_only_one_lane(
bbox: BBox,
best_lane: _TextLane,
lanes: list[_TextLane],
tolerance: float,
) -> bool:
"""保留旧内部入口;独立调用仍自行排序,输出与预排序版完全一致。"""

return _fits_only_one_lane_ordered(
bbox,
best_lane,
sorted(lanes, key=lambda lane: lane.left),
tolerance,
)


def _expand_nested_lane_intervals_from_members(
lanes: list[_TextLane],
tolerance: float,
Expand Down Expand Up @@ -733,6 +750,10 @@ def _reattach_cross_lane_short_tails(
median_height: float,
) -> None:
"""按纵向事件复用各栏最近前序行;特殊输入交给原逐次扫描实现。"""
# 只有一个栏带时不存在“跨栏”归属,直接跳过输入校验和前序扫描;
# 这不改变成员、顺序或边界,也能让常见单栏页面避开整页重复遍历。
if len(lanes) < 2:
return
pending = []
seen_lines, seen_indices = set(), set()
for lane_index, lane in enumerate(lanes):
Expand Down
72 changes: 45 additions & 27 deletions src/docvortex/analyzers/native/pdf/table_annotations.py
Original file line number Diff line number Diff line change
Expand Up @@ -162,40 +162,58 @@ def prepared_marker_line(self, line: _LineItem, page_size: tuple[float, float],
return self.marker_prepared.prepare(line, page_size, angle)


def _plain_marker_bbox(value) -> bool:
"""按 marker 快路径的既有规则校验字符 bbox 形状和数值范围。"""
if type(value) not in (tuple, list, CharBbox):
return False
raw = value.bbox if type(value) is CharBbox else value
if type(raw) not in (tuple, list) or len(raw) != 4:
return False
for number in raw:
if type(number) is float:
if not math.isfinite(number):
return False
elif type(number) is int:
if not -(2**53) <= number <= 2**53:
return False
else:
return False
return True


def _prepare_marker_line_context(lines):
"""同一候选上下文只校验一次 marker 输入,并冻结来源行索引。"""
from ...._compute_backend import get_native

if get_native() is None:
return False, {}
source_lines: dict[int, list[_LineItem]] = {}
safe = all(
type(line) is _LineItem
and type(line.source_index) is int
and type(line.text) is str
and type(line.chars) is list
and all(
type(char) is dict
and (char.get("char") is None or type(char.get("char")) is str)
and (
char.get("bbox") is None
or (
type(char.get("bbox")) in (tuple, list, CharBbox)
and len(char["bbox"].bbox if type(char["bbox"]) is CharBbox else char["bbox"]) == 4
and all(
(type(value) is float and math.isfinite(value)) or (type(value) is int and -(2**53) <= value <= 2**53)
for value in char["bbox"]
)
)
)
for char in line.chars
)
for line in lines
)
if safe:
for line in lines:
source_lines.setdefault(line.source_index, []).append(line)
return safe, source_lines
# 直接单次遍历替代多层生成式,避免同一 char/bbox 字典键被重复读取;
# 任一特殊输入立即返回 False,保持原参考路径整体回退的语义。
for line in lines:
if (
type(line) is not _LineItem
or type(line.source_index) is not int
or type(line.text) is not str
or type(line.chars) is not list
):
return False, {}
plain_chars = True
for char in line.chars:
if type(char) is not dict:
plain_chars = False
break
text = char.get("char")
if text is not None and type(text) is not str:
plain_chars = False
break
if not _plain_marker_bbox(char.get("bbox")) and char.get("bbox") is not None:
plain_chars = False
break
if not plain_chars:
return False, {}
source_lines.setdefault(line.source_index, []).append(line)
return True, source_lines


def _prepare_table_core_rows(
Expand Down
29 changes: 16 additions & 13 deletions src/docvortex/analyzers/native/pdf/table_detection.py
Original file line number Diff line number Diff line change
Expand Up @@ -77,20 +77,22 @@ def _detect_table_candidates(
*local_excluded_bboxes,
*[_rotate_bbox_to_upright(bbox, source.page_size, angle) for bbox in source.form_bboxes],
]
rule_candidates.extend(
_build_rule_table_candidates(
rows,
angle_lines,
source.page_size,
angle,
median_height,
local_axis_lines,
path_infos=source.path_infos,
excluded_bboxes=local_excluded_bboxes,
caption_candidates=caption_candidates,
defer_materialization=True,
)
angle_rule_candidates = _build_rule_table_candidates(
rows,
angle_lines,
source.page_size,
angle,
median_height,
local_axis_lines,
path_infos=source.path_infos,
excluded_bboxes=local_excluded_bboxes,
caption_candidates=caption_candidates,
defer_materialization=True,
)
rule_candidates.extend(angle_rule_candidates)
# 规则候选与闭合网格消费同一批横线/竖轨;复用上下文冻结的分量,
# 避免 pollutant 这类多表页面连续两次重建相同连通关系。
shared_grid_components = angle_rule_candidates[0].context.grid_components if angle_rule_candidates else None
rule_candidates.extend(
_build_closed_rule_grid_candidates(
rows,
Expand All @@ -101,6 +103,7 @@ def _detect_table_candidates(
local_axis_lines,
local_closed_grid_excluded_bboxes,
caption_candidates,
grid_components=shared_grid_components,
)
)
merged_rule_candidates = [
Expand Down
Loading
Loading