Skip to content

feat: Chinese retrieval enhancements - pluggable BM25 tokenizers, per-doc diversity cap, Gradio Web UI - #2

Open
hulinming wants to merge 3 commits into
IvenKooLab:mainfrom
hulinming:dev/zh-enhance
Open

hulinming wants to merge 3 commits into
IvenKooLab:mainfrom
hulinming:dev/zh-enhance

Conversation

@hulinming

Copy link
Copy Markdown
Contributor

Summary

A batch of enhancements focused on Chinese-language retrieval and usability. All changes are backward compatible — the original dependency-free single-char CJK tokenizer stays the default, and every addition ships as an optional extra.

1. Pluggable BM25 tokenizers with jieba Chinese word segmentation

  • src/loci/tokenizers.py: tiny registry (default / jieba); jieba is imported lazily and falls back to default with a one-time install hint (fail-open).
  • Config: [bm25] tokenizer = "default" | "jieba"; unknown names warn and fall back to default.
  • Optional extra: pip install 'loci-rag[zh]' — no new runtime dependency.
  • Effect: the default tokenizer splits Chinese into single characters, so phrase-level BM25 recall is weak (A/B on a real index: a distractor doc outranked the correct one 0.416 vs 0.325); with jieba the correct doc scores 0.595 and the distractor 0.000.

2. Per-document diversity cap

  • Retriever keeps at most max_per_doc chunks (default 2) per source document in one result list, applied after RRF fusion/rerank/filters so surplus chunks give way to other relevant documents.
  • Config: [retrieval] max_per_doc (0 = no cap, preserves the original behavior).
  • Effect: top-5 results spread across 3 documents instead of being dominated by one long file.

3. Gradio Web UI (optional)

  • New loci webui command: chat pane + per-answer evidence panel that reuses build() / make_retriever() (same pipeline as the CLI and HTTP API); evidence shows document name > section.
  • Optional extra: pip install 'loci-rag[ui]'.

4. Smaller items

  • SYSTEM_PROMPT rewritten as strict grounding rules: excerpts only, explicit "not in KB" answer instead of guessing, bullet structure, mandatory Sources: citations.
  • search / ask / bench accept --top_k as a long-form alias of -k (same args.k, no behavior change).
  • Missing optional deps (pdf / docx / ocr / jieba) now warn once per run with the exact pip install command instead of skipping files silently.
  • .gitignore: added .venv/.

Tests

  • 207 passed, 1 skipped — all offline (FakeEmbedder; gradio/jieba tests are importorskip-guarded), +24 new tests in this branch.

The original tokenizer splits CJK into single characters, so phrase-level
keyword recall fails for Chinese: query 检索 matches a document that merely
contains 检 and 索 inside different words, and that decoy can outrank the
genuinely relevant chunk.

- new loci.tokenizers registry: "default" (original behavior, no deps) and
  "jieba" (word-level segmentation, lazy import, fails open to default when
  the optional package is missing)
- BM25 accepts an injected tokenizer; module-level tokenize() kept as a
  backward-compatible alias (conftest FakeEmbedder relies on it)
- [bm25] tokenizer config section wired through Retriever; unknown names
  warn and fall back to default at load time
- bench gains --tokenizer so A/B runs need no config edits

A/B on a decoy corpus (query 检索; doc b contains 检+索 in other words):
default ranks decoy 0.416 > correct 0.325; jieba ranks correct 0.595 and
zeroes the decoy. 10 new tests in tests/test_tokenizers.py.
…extra)

- src/loci/webui.py: chat pane + per-answer evidence panel reusing
  build()/make_retriever() from the CLI pipeline; ask_fn keeps the last
  four turns as history; format_evidence() renders source > section
- cli: new `webui` subcommand with --host/--port overrides
- config: [webui] section (host 127.0.0.1, port 7860)
- pyproject: ui = ["gradio>=4"] optional extra (no new runtime deps)
- tests/test_webui.py: 10 offline tests (evidence formatting, ask
  callback, history pass-through, app construction, config defaults)
…g prompt, --top_k alias, missing-extra guidance

- retriever: max_per_doc (default 2) caps chunks per source document in
  one result list, applied after RRF fusion/rerank so surplus chunks give
  way to other docs; [retrieval] max_per_doc config (0 = no cap)
- retriever: SYSTEM_PROMPT rewritten as 4 strict rules — excerpts only,
  explicit "not in KB" answer, bullet structure, mandatory Sources line
- cli: search/ask/bench accept -k/--top_k alias (dest=k keeps compat)
- webui: evidence panel shows document name > section (full path stays in
  the answer citations)
- loaders/tokenizers: missing pdf/docx/ocr/jieba extras now warn once per
  run with the exact pip install command instead of skipping silently
- tests: +14 (diversity cap, prompt rules, --top_k wiring, dep warnings)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant