Skip to content

Sync OpenCC dictionaries, configs, and upstream testcases to version 1.4.1 - #34

Open
frankslin wants to merge 2 commits into
yichen0831:masterfrom
nk2028:master
Open

Sync OpenCC dictionaries, configs, and upstream testcases to version 1.4.1#34
frankslin wants to merge 2 commits into
yichen0831:masterfrom
nk2028:master

Conversation

@frankslin

Copy link
Copy Markdown
  • Replace bundled OpenCC dictionary sources with the current upstream text dictionaries.
  • Replace bundled conversion configs with upstream configs, converting .ocd2 references to .txt so they work with this pure-Python loader.
  • Generate missing reverse dictionaries for HK, TW, and JP variants from the updated upstream source dictionaries.
  • Remove obsolete split Taiwan phrase dictionaries: TWPhrasesIT.txt, TWPhrasesName.txt, and TWPhrasesOther.txt.
  • Add a data-driven unittest that runs the upstream testcase corpus against all expected conversion modes.
  • Add upstream testcase data into opencc/testcases/ so imported upstream assets live under the package tree.
  • Teach dictionary loading to ignore upstream comment/blank lines, and allow ASCII hyphenated terms to match dictionary entries.
  • Fix the existing unittest import path so tests can be run from the repository root.

frankslin added 2 commits May 8, 2026 16:28
- Replace bundled OpenCC dictionary sources with the current upstream text dictionaries.
- Replace bundled conversion configs with upstream configs, converting `.ocd2` references to `.txt` so they work with this pure-Python loader.
- Generate missing reverse dictionaries for HK, TW, and JP variants from the updated upstream source dictionaries.
- Remove obsolete split Taiwan phrase dictionaries: `TWPhrasesIT.txt`, `TWPhrasesName.txt`, and `TWPhrasesOther.txt`.
- Add a data-driven unittest that runs the upstream testcase corpus against all expected conversion modes.
- Add upstream testcase data into `opencc/testcases/` so imported upstream assets live under the package tree.
- Teach dictionary loading to ignore upstream comment/blank lines, and allow ASCII hyphenated terms to match dictionary entries.
- Fix the existing unittest import path so tests can be run from the repository root.
- Refresh bundled source dictionaries from upstream ver.1.4.1 and
  regenerate reverse dictionaries via helper/reverse.py.
- Restructure Japanese configs for the upstream Shinjitai refactor:
  t2jp now uses a generated JPShinjitaiCharactersRev, jp2t uses the
  JPShinjitaiPhrases + JPShinjitaiCharacters group, and the obsolete
  JPVariants/JPVariantsRev dictionaries are removed.
- Add Hong Kong phrase modes s2hkp and hk2sp with bundled HKPhrases
  and HKPhrasesRev dictionaries.
- Fix helper/reverse.py to skip comment/blank lines so upstream's new
  "# Format:" header is not reversed into spurious dictionary entries.
- Regenerate opencc/testcases/testcases.json from the ver.1.4.1 corpus
  (192 cases, 441 assertions across all 16 modes), keeping only the
  supported-mode conversions the pure-Python port reproduces exactly.
- Update packaging: bump version to 0.1.8, use globs in package_data,
  include *.json in MANIFEST.in, and document all 16 modes in README.
@frankslin frankslin changed the title Sync OpenCC dictionaries, configs, and upstream testcases Sync OpenCC dictionaries, configs, and upstream testcases to version 1.4.1 Jul 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant