Entity matching system for Japanese insurance customer data. Resolves duplicate records across policies, agencies, and channels.
Read ARCHITECTURE.md (one page) and AGENTS.md (rules + "how to add a view/endpoint/decision" checklists). One command for status:
python -m api.cli audit repo # ruff, mypy, tsc, npm test, pytest (+ E2E opt-in)The original root-level app layer (nayose_cli.py) was retired and the engine
migrated to src/api/infrastructure/runtime/engine.py; all behavior lives in src/api/, frontend/,
src/data_generation/.
Reproducible environments: open this repo in VSCode Dev Containers (
.devcontainer/) for a configured dev box, or pull the tagged image from GHCR. Git/CI/Docker/cloud workflow:docs/CLOUD.md; ops runbook:docs/OPERATIONS.md; GPU training:docs/REMOTE_GPU.md. Evidence-first matching and the Japanese/English UI:docs/MATCHING_UI.md. 5-minute stakeholder demo:docs/DEMO.md.
The generator writes the visible snapshots (reference_records,
incoming_queries) plus core hidden truth tables and clean policy lifecycle
tables under data/ground_truth/.
source .venv/bin/activate
uv sync --extra matcher --extra neural --extra performance --extra serve --extra dev
python -m api.cli data generate --profile smoke --data-dir data --seed 42
# profiles: ~340 / ~800 / ~6,800 / ~22,700 / ~113,500 canonical entities
# one-time downloads, cached in data/raw/: Nemotron persona names, Japan Post ken_all, EDINETFor a large canonical layer, stream clean entities without retaining sabotaged records, policies, and lineage in memory:
python -m api.cli data generate --clean-only --n-parties 2000000 \
--chunk-size 50000 --seed 42 --data-dir data
# writes data/clean_generations/clean_<timestamp>_<hash>/ and updates data/clean_currentThe full generation path remains capped at 2,000,000 base entities because it
materializes all observations, policy lifecycle tables, and field lineage before
saving. The clean-only path is chunked and is the implementation route for a 2M
canonical entity layer; target-runner capacity sign-off is still required, and it
does not replace data/current.
# Supported headless path: runs the locked F2LLM-v2-330M pipeline on Apple MPS,
# validates the bundle, and materializes the exact/ANN artifacts.
python -m api.cli model train --data-dir data/current --models-dir models \
--seed 42 --device mps --model-family f2llm_330m --pair-model lightgbmnayose model train runs the isolated api.training.worker, which delegates to the modular src/api/training/ application, then writes a model bundle
(portable indexes + target-platform ANN + run_manifest) into its output
directory. NayoseEngine.load fingerprint-checks the bundle against
the
reference data and refuses stale combinations:
from api.infrastructure.runtime.engine import load_engine
engine = load_engine() # models/ + data/ (after a fresh nayose run)
result = engine.recommend({"name_kanji": "山田太郎", "phone": "090-1234-5678"})
print(result.action) # AUTO_LINK
print(result.top_nayose_id) # N000012345Headless training and serving use the same immutable bundle contract:
# then open http://127.0.0.1:8899
NAYOSE_DASH_PORT=8899 NAYOSE_WARM_EMBED=1 python -m api.main├── src/ # Installable Python packages
│ ├── api/cli.py # canonical data/model/audit/benchmark CLI facade
│ ├── api/commands/ # focused command implementations
│ ├── api/http/routers/ # typed FastAPI endpoint adapters
│ ├── api/application/ # matching, search, query, and admin use cases
│ ├── api/domain/ # normalization, evidence, and identity rules
│ ├── api/infrastructure/ # bundle, storage, retrieval, and runtime adapters
│ ├── api/training/ # isolated modular training application
│ └── data_generation/ # sources, generator, corruption, truth, and workflow
│
├── data/ # Runtime data products (gitignored)
│ ├── raw/ # immutable reference data (ken_all, personas, EDINET)
│ ├── generations/gen_<ts>_<hash>/ # one immutable dataset product per run
│ │ ├── reference_records.parquet / incoming_queries.parquet / agencies.parquet
│ │ └── ground_truth/ # identity + clean policy truth tables
│ ├── current -> generations/gen_<ts>_<hash> # promoted generation (consumer-facing)
│ │ ├── record_truth.parquet
│ │ ├── party_truth.parquet
│ │ ├── relation_truth.parquet
│ │ ├── query_truth.parquet
│ │ ├── evaluation_truth.parquet
│ │ ├── field_lineage.parquet # per-field provenance audit (1.9M rows @100k)
│ │ ├── policy_truth.parquet / policy_term_truth.parquet
│ │ ├── policy_event_truth.parquet / policy_party_truth.parquet
│ │ └── policy_risk_truth.parquet
│ └── operational.sqlite3 # SQLite WAL jobs, decisions, audit, manual staging
│
└── models/ # Trained artifacts (gitignored)
├── runs/run_<ts>_<hash>/ # immutable bundle + run_manifest.json
├── current -> runs/... # validated serving alias
└── registry.json # append-only run metadata
Static HLD/LLD diagrams are generated from the Python and TypeScript import graphs:
python scripts/generate_architecture_diagrams.py; see
docs/ARCHITECTURE_DIAGRAMS.md.
| Component | Responsibility |
|---|---|
InferenceConfig |
Immutable serving settings |
CandidateIndex |
Candidate retrieval |
NayoseEngine |
Artifact loading, validation, and recommendation |
normalize_*() |
Field normalization |
NAYOSE_MODEL_FAMILY defaults to off, the reproducible CPU/SVD baseline. Neural families are explicit opt-ins:
| Family | Embedding | Notes |
|---|---|---|
off (default) |
— | Deterministic SVD baseline |
f2llm_330m |
codefuse-ai/F2LLM-v2-330M |
Locked production profile: Apple MPS, FP16, batch 32, one process |
f2llm |
codefuse-ai/F2LLM-v2-330M |
Compatibility alias for the locked 330M profile |
f2llm_0.6b |
codefuse-ai/F2LLM-v2-0.6B |
Experimental MPS-only alternative, no fallback |
ruri |
Ruri encoder/reranker path | Optional A/B entry |
qwen |
Qwen embedding/reranker path | Optional |
The production neural path is MPS-only and fail-closed: unavailable MPS, missing
sentence-transformers, model-load errors, and unsupported fallback settings abort the
run. The locked F2LLM-v2-330M settings are FP16, batch 32, one process, dense retrieval
enabled, and PYTORCH_ENABLE_MPS_FALLBACK=0. Record the selected family, backend,
device, and degraded status in the run manifest before comparing runs.
Optimizations for MPS + time complexity:
- embedding encode on MPS in fp16, batched (
NAYOSE_NEURAL_BATCH); - reference embeddings cached on disk keyed by (data fingerprint, model id) — encode once per generation, later runs are O(reload);
- usearch (HNSW) ANN index for dense retrieval instead of brute-force;
- exact blocks + character sparse retrieval remain as high-precision anchors.
One-time cold-encode cost is paid once per immutable generation and cached by
(data_fingerprint, model_id). Usef2llm_330mfor the production candidate path; benchmark the actual device before selecting batch sizes.
Reading (kana) is the identity; observed kanji may diverge realistically:
- 同音異字 (homophone kanji): same kana reading, different kanji, drawn from
the REAL name corpus (e.g. 前田道枝 → 前田美知恵, both マエダ ミチエ). Pools are
built data-driven from the persona cache (
build_homophone_map), no lexicon. - kana forms: same reading rendered hiragana or half-width katakana.
- digit transposes / email typos: the classic typist mistakes.
- Systematic (vintage/system) field availability: legacy records never hold email (keep=0.0), paper rarely do (0.45), web always -- missingness is channel/system-driven, not uniformly random.
- Source-document bundles: a record typed from a channel's form carries the channel's characteristic disorders together (web: email+digit typos; agent_pc: homophone+digit; paper_ocr: OCR+kana-form; legacy: era+width+no email).
Tuning: every rate/table lives in NoiseConfig in src/data_generation/config.py
(event_rates, field_availability, channel_events, band_scale,
missing_rates) and the twin rate via NAYOSE_TWIN_RATE. Truth stays
canonical (party_truth) and lineage-exact; records are explainable derivations.
The synthetic dirty data is driven from truth, not random noise:
every observed record is a logical, explainable derivation of its canonical
entity (name renames, moves, re-contacts, pre-update snapshots, OCR scans,
legacy era formats), and the truth tables + field_lineage remain the exact
answer for each value.
Hardness levers (all in src/data_generation/corruption.py / generator.py):
- Coherent events:
rename_stale,move_stale,contact_change,full_snapshot(pre-move/pre-re-contact record using the entity's old values),contact_omission(phone+email missing together),paper_ocr,legacy_format,doc_damage— each internally consistent and labelled. - Twin entities (
NAYOSE_TWIN_RATE, default 0.15): additional customers with the SAME name/kana/romaji but different dob/address/contacts —hard_case=twin, exact-name anchors become ambiguous. - Channel attribution:
NAYOSE_EXACT_MODE = hybrid | fuzzy_only | exact_only— run the matcher with exact blocks disabled to measure how much retrieval relies on fuzzy (sparse+dense) evidence. - Hardness telemetry: the run prints e.g.
HARDNESS: 37% of TRUE pairs carry ZERO exact-block anchors and must be found fuzzilyand the generator prints name-duplication/twin stats; the derive target is retrieval-hard data where dense embeddings visibly matter.
Encoding decision: index the normalized observed record (distribution
matches serve-time dirty queries; party_truth is hidden at deploy time).
F2LLM candidate discovery also materializes separate record, name, address, and
search-only alias views; aliases retain source-record provenance and never
rewrite entity profiles. Record candidates are fused with exact and sparse
channels before structured pair verification, so dense similarity never becomes
an identity decision by itself.
Runtime optimization settings are defined only in the canonical
src/nayose/nayose.toml. Deployment environment overrides remain available for
secrets, paths, and approved operational controls, but they do not change
embedding backend, embedding batch size, embedding worker count, or MLX
sequence/padding behavior.
| Group | Env | Default | Meaning |
|---|---|---|---|
| Data gen | NAYOSE_SEED |
42 | RNG seed (reproducibility) |
NAYOSE_PROFILE |
generate_only | 300/700/6,000/20,000/100,000 base entities | |
NAYOSE_N_PARTIES |
0 | full-generation base entity override; hard cap 2,000,000 | |
NAYOSE_DIFFICULTY |
realistic | debug/realistic/challenge/severe (dirtiness rates) | |
| Retrieval | NAYOSE_EXACT_BLOCK_CAP |
80 | block-exhaustive cap (skips oversized blocks) |
NAYOSE_DENSE_VIEWS |
1 | materialize field-aware F2LLM name/address/alias artifacts | |
NAYOSE_SPARSE_TOPK / NAYOSE_DENSE_TOPK |
24 / 16 | k per retrieval channel | |
NAYOSE_CANDIDATE_CAP |
80 | max candidate records per query | |
NAYOSE_SPARSE_BATCH / NAYOSE_NEURAL_BATCH |
512 / 32 | legacy command inputs; embedding batch size is TOML-only and 32 is frozen for F2LLM | |
| Embeddings | NAYOSE_EMBEDDING_MODEL |
F2LLM-v2-330M | locked production dense encoder |
NAYOSE_DEVICE |
mps | locked PyTorch/MPS compatibility device | |
PYTORCH_ENABLE_MPS_FALLBACK |
0 | unsupported MPS ops fail instead of using CPU | |
NAYOSE_USE_DENSE |
1 | dense channel on/off; bundle artifacts remain strict | |
NAYOSE_CPU_THREADS |
auto | per-process numerical thread budget | |
NAYOSE_PROCESS_WORKERS |
1 | process count used to divide non-embedding CPU threads | |
NAYOSE_INDEX_BACKEND |
polars | bulk index materialization backend | |
NAYOSE_SVD_DIM |
32 | SVD dimensions for the off model |
|
| Pair model | NAYOSE_PAIR_MODEL |
auto | auto/lightgbm/sklearn |
NAYOSE_NEG_PER_POS |
6 | negatives per positive in training sample | |
NAYOSE_MODEL_THREADS |
auto | bounded model-training threads | |
| Calibration/decision | NAYOSE_PAIR_AUTO_PRECISION |
0.985 | Wilson lower-bound target (train-set-proof) |
NAYOSE_PAIR_REVIEW_PRECISION |
0.80 | review threshold target | |
NAYOSE_ENTITY_AUTO_PRECISION |
0.97 | entity auto precision target | |
NAYOSE_ENTITY_MARGIN |
0.06 | margin over second entity | |
NAYOSE_ENTITY_OVERRIDE |
0 | force manifest thresholds off |
For Apple Silicon MLX can be selected only in the canonical TOML:
[embedding]
backend = "mlx_gpu" # mlx_cpu is the CPU fallback; pytorch_mps is the reference
[embedding.optimization]
max_length = 0 # model maximum when zero
padding_strategy = "longest"The MLX extra is optional: uv pip install -e '.[mlx]'. MLX is fail-closed
when unavailable or when the model architecture is unsupported; it never
silently falls back to PyTorch.
LightGBM internals (n_estimators=260, lr=0.035, num_leaves=31, …) live in the modular pair-training stage; the sklearn HistGB backend is the MPS-serving alternative.
- Unseen test set: edges/queries with
pair_split/query_split == testare never used for training or threshold calibration (also reported per-split). - Every run saves
pair_confusion_test.parquet(TP/FP/TN/FN, precision, recall, specificity, F1, AUC @ auto threshold) andquery_confusion.parquet(predicted action × expected-existence per scenario) into the bundle, and prints both with the run. - Candidate (retrieval) recall is reported separately from classifier precision exactly so retrieval ease is not mistaken for matching accuracy.
A no-build modern frontend (frontend/) + thin API (src/api/main.py), served
locally by one process. DuckDB executes read-only SQL over the run bundle
parquets (single local process; swap the data adapter to DuckDB-WASM later).
The interface uses progressive disclosure and shows verdict uncertainty with
lineage in the flow. Routes are URL-addressable (#/entity/:id,
#/compare/:l/:r), the palette is keyboard accessible, and reduced-motion and
AA contrast requirements are covered by the frontend checks.
NAYOSE_MODELS_DIR=models NAYOSE_DATA_DIR=data/current \
.venv/bin/python -m api.main # → http://127.0.0.1:8000Views: 検索(既定画面)・照合テスト (live /api/recommend) · レビュー待機
(件数・確率フィルタ・既読保存)・レコード比較(パンくず・エンティティ連携)・
混同行列(ヒート・内訳ドリル)・指標(バー付き)・エンティティ探索(検索・
サーバー側ページング・代表情報・出所関係)・管理(複数世代のデータ閲覧、
全体/列フィルタ、固定列、ページング)。
Frontend boundaries: frontend/ts/schema.ts + core.ts (pure schema/core) ·
port.ts (typed capability and storage interfaces) · infra.ts (ApiAdapter +
cache, LocalStorageStore) · feature modules (views depend only on ports; the
admin dataset grid is isolated in dataset-grid.ts) ·
store.ts (observable state with reversible decisions) · app.ts assembly
(boot() injects port/store/runs; no import side effects).
API (typed domain endpoints; raw SQL only behind NAYOSE_DEBUG_SQL=1):
/api/runs · /api/entities?q= · /api/entities/explorer ·
/api/entity/{id} (lineage) · /api/entity/{id}/graph ·
/api/review?min_p=&offset= · /api/metrics · /api/confusion ·
/api/compare · /api/recommend (fingerprint-guarded) · /api/admin/generations
· /api/admin/generations/{id}/summary · /api/admin/generations/{id}/records
(global/column filters, bounded pages, optional keyset after cursor) · /api/stats
(per-endpoint hits+ms). Caching: server-side per-run LRU + client adapter memo;
field lineage remains in the immutable generation's hidden field_lineage table.
Cross-run diff (#/compare-runs): pick run A/B, ポート.metricsFor(run) →
side-by-side KPI table with Δ (green/red deltas, tabular numerals).
Decision persistence (server-side): review 既読 is a Command in the store,
persisted in the operational SQLite store via GET/PUT /api/decisions
(debounced) — decisions survive across machines/sessions; localStorage remains
the offline fallback; 取り消し undoes via the command history.
E2E (real browser, opt-in): NAYOSE_E2E=1 NAYOSE_E2E_MODELS=… NAYOSE_E2E_DATA=… npm --prefix frontend run test:e2e — starts the server, drives Chromium
(Playwright), asserts search→entity→review persistence→cross-run diff and the
⌘K palette flow.
Testing policy (hermetic, deterministic, enforced):
- TS unit:
frontend/test/core.test.mjs(pure core, node:test). - TS functional:
frontend/test/view.test.mjs— boots the app with a MockPort in jsdom; asserts real flows (default route, search, palette keyboard, review check+undo, confusion drill). No network. - Python: API functional tests over a fake bundle + cross-language schema
contract test (schema.ts ↔ REQUIRED_COLUMNS) +
tests/test_frontend_gate.pyrunsnpm test(skips when node/npm absent).
Frontend is strict TypeScript compiled to ESM (tsc), no bundler:
(cd frontend && npm install) # once (typescript)
(cd frontend && npm run typecheck) # strict type gate
(cd frontend && npm run build) # ts/ -> js/ (committed, served directly)Contracts are Rust-disciplined: frontend/ts/schema.ts mirrors
src/api/http/contracts.py REQUIRED_COLUMNS, and tests/test_schema_contract.py fails CI
if they drift; the data adapter validates shapes at runtime before render.
Never hand-edit frontend/js/ (generated).
| View | Content |
|---|---|
| F1 指標 | candidate recall / B³ F1 / top-1 / MRR / review / no-match |
| F2 混同行列 | pair verifier (unseen test) + query-decision confusion |
| F3 レビュー待機 | uncertain twin/edge queue |
| F4 エンティティ | cluster cards: canonical + members + field-lineage summary |
| F5 検索 | debounced fuzzy search over names/phone/postal, capped results |
Design: task-first progressive disclosure, light theme + Japanese labels
(Noto Sans JP), semantic HTML (role landmarks, aria-live, AA contrast,
focus-visible, container queries, reduced-motion).
Front-end checklist mapping (thedaviddias/front-end-checklist): ticked where applicable (HTML semantics, CSS responsiveness, a11y, performance: no chart library, schema-validated at render); single-file inline CSS/JS is a documented exception for a local offline artifact; web-deployment-only rules (HTTP caching, CDN/SRI, SEO crawlers, PWA) are not-applicable to a file — mapping maintained in AGENTS.md.
A/B on regenerated 100k-hard data (113,499 parties, 24,522 same-name twin rows, 249,201 records; fuzzy-only retrieval, ruri-v3-130m, exact blocks off):
| Metric | sparse+dense | sparse-only |
|---|---|---|
| Candidate recall | 99.34% | 99.12% |
| Top-1 identity (existing) | 90.9% | 79.8% |
| MRR | 0.933 | 0.809 |
| No-match rate | 16.9% | 24.1% |
| B³ F1 | 0.971 | 0.970 |
| Embedding encode | 47 min once (cached) | 0 |
The dense-vs-sparse identity gap grows with scale+hardness (smoke +2.8pts, mac +5.6, 100k +11.1): sparse char-ngrams saturate retrieval on small data but cannot disambiguate same-name twins at scale. Dense embeddings are the ranking/robustness layer; keep them on by default.
src/api/ installable serving, application, domain, and CLI package
src/data_generation/ installable data-generation package
frontend/ no-build ESM app (index.html, styles.css, js/app.js)
tests/ unit, integration, and architecture/contract tests
ci/ quality and resource-budget configuration
docs/ architecture, operations, and deployment guides
data/, models/ gitignored runtime products; never source code
nayose model train writes each train/evaluate into an immutable model bundle and promotes it via a symlink:
models/
├── registry.json # run_id → bundle, config/data hashes, backends, stage metrics
├── runs/run_<ts>_<hash>/ # frames + models + sparse/optional inverted index + embeddings + manifest
└── current -> runs/run_<ts>_<hash> # promoted bundle (consumer-facing)
NayoseEngine.load(model_dir="models/current", data_dir="data/current"):
- fingerprint-checks bundle vs reference (refuses stale data↔model pairings);
- loads sparse TF-IDF + F2LLM dense embeddings with usearch HNSW;
- loads the pair verifier — native LightGBM text booster when available, else joblib;
- serves
recommend()= normalize → blocks ∪ sparse ∪ dense retrieval → calibrated pair probabilities → entity decision.
MPS note: joblib-restored LightGBM and sentence-transformers encoding deadlock in one process on Apple Silicon. Serve the pair model via the native text booster (default when trained with LightGBM), or train with
NAYOSE_PAIR_MODEL=sklearnfor a pure-sklearn serving bundle. Both are recorded inregistry.json.
| Metric | Value |
|---|---|
| Bundle load | ~7.3 s (374,677-row MPS benchmark, guard passes) |
| Inference (warm) | ~0.032 s/query (20-iteration p50) |
| Sparse retrieval | deterministic top-k; 64 highest-weight query features by default |
| Training (smoke, f2llm+MPS) | ~48 s end-to-end (embeddings cached) |
| B³ F1 (smoke) | 96.8–96.9 %; top-1 identity 96.8 % (f2llm) vs 94.7 % (off) |
numpy pandas scipy scikit-learn rapidfuzz lightgbm joblib
pykakasi phonenumbers jaconv pyarrow faker(optional legacy) datasets
The generator logic lives in src/data_generation/ (sources / corruption / ground truth / generator;
data-driven config in src/data_generation/config.py), with nayose data generate as
its thin CLI.
uv pip install pytest && python -m pytest tests/ -q| Source | Use | Licence / credit |
|---|---|---|
| NVIDIA Nemotron-Personas-Japan | names + sex/age/prefecture/occupation (demographic skew) | CC-BY 4.0 |
| Japan Post ken_all | ~119k postal towns, all 47 prefectures | Japan Post postal-code data |
| EDINET edinet2023 | 4,253 real company legal names | EDINET (FSA) |
- Correlated events (rename/move/contact stale, paper-OCR bursts, legacy format incl. era dates + half-width kana, correlated missing) scaled by quality band × agency factors; every value carries a lineage label.
- Orthographic consistency: kana/romaji are derived from the kanji, so names never disagree across spellings (82% of field cells stay clean).
- Benchmark validated end-to-end: smoke profile ->
nayose model trainB3-F1 ≈ 96.5%, candidate recall ≈ 99.8% (modeloff, CPU).