Repository navigation
Conversation
The original tokenizer splits CJK into single characters, so phrase-level keyword recall fails for Chinese: query 检索 matches a document that merely contains 检 and 索 inside different words, and that decoy can outrank the genuinely relevant chunk. - new loci.tokenizers registry: "default" (original behavior, no deps) and "jieba" (word-level segmentation, lazy import, fails open to default when the optional package is missing) - BM25 accepts an injected tokenizer; module-level tokenize() kept as a backward-compatible alias (conftest FakeEmbedder relies on it) - [bm25] tokenizer config section wired through Retriever; unknown names warn and fall back to default at load time - bench gains --tokenizer so A/B runs need no config edits A/B on a decoy corpus (query 检索; doc b contains 检+索 in other words): default ranks decoy 0.416 > correct 0.325; jieba ranks correct 0.595 and zeroes the decoy. 10 new tests in tests/test_tokenizers.py.
…extra) - src/loci/webui.py: chat pane + per-answer evidence panel reusing build()/make_retriever() from the CLI pipeline; ask_fn keeps the last four turns as history; format_evidence() renders source > section - cli: new `webui` subcommand with --host/--port overrides - config: [webui] section (host 127.0.0.1, port 7860) - pyproject: ui = ["gradio>=4"] optional extra (no new runtime deps) - tests/test_webui.py: 10 offline tests (evidence formatting, ask callback, history pass-through, app construction, config defaults)
…g prompt, --top_k alias, missing-extra guidance - retriever: max_per_doc (default 2) caps chunks per source document in one result list, applied after RRF fusion/rerank so surplus chunks give way to other docs; [retrieval] max_per_doc config (0 = no cap) - retriever: SYSTEM_PROMPT rewritten as 4 strict rules — excerpts only, explicit "not in KB" answer, bullet structure, mandatory Sources line - cli: search/ask/bench accept -k/--top_k alias (dest=k keeps compat) - webui: evidence panel shows document name > section (full path stays in the answer citations) - loaders/tokenizers: missing pdf/docx/ocr/jieba extras now warn once per run with the exact pip install command instead of skipping silently - tests: +14 (diversity cap, prompt rules, --top_k wiring, dep warnings)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A batch of enhancements focused on Chinese-language retrieval and usability. All changes are backward compatible — the original dependency-free single-char CJK tokenizer stays the default, and every addition ships as an optional extra.
1. Pluggable BM25 tokenizers with jieba Chinese word segmentation
src/loci/tokenizers.py: tiny registry (default/jieba); jieba is imported lazily and falls back todefaultwith a one-time install hint (fail-open).[bm25] tokenizer = "default" | "jieba"; unknown names warn and fall back todefault.pip install 'loci-rag[zh]'— no new runtime dependency.2. Per-document diversity cap
Retrieverkeeps at mostmax_per_docchunks (default 2) per source document in one result list, applied after RRF fusion/rerank/filters so surplus chunks give way to other relevant documents.[retrieval] max_per_doc(0 = no cap, preserves the original behavior).3. Gradio Web UI (optional)
loci webuicommand: chat pane + per-answer evidence panel that reusesbuild()/make_retriever()(same pipeline as the CLI and HTTP API); evidence shows document name > section.pip install 'loci-rag[ui]'.4. Smaller items
SYSTEM_PROMPTrewritten as strict grounding rules: excerpts only, explicit "not in KB" answer instead of guessing, bullet structure, mandatorySources:citations.search/ask/benchaccept--top_kas a long-form alias of-k(sameargs.k, no behavior change).pip installcommand instead of skipping files silently..gitignore: added.venv/.Tests
importorskip-guarded), +24 new tests in this branch.