Skip to content

Tamil TN Cardinal Semiotic Class - #449

Open
surendran-246 wants to merge 37 commits into
NVIDIA:staging/tamil_tn_v1from
surendran-246:feat-ta-cardinal
Open

Tamil TN Cardinal Semiotic Class#449
surendran-246 wants to merge 37 commits into
NVIDIA:staging/tamil_tn_v1from
surendran-246:feat-ta-cardinal

Conversation

@surendran-246

@surendran-246 surendran-246 commented Jul 7, 2026

Copy link
Copy Markdown

What does this PR do ?

Add a one line overview of what this PR aims to accomplish.

Before your PR is "Ready for review"

Pre checks:

  • Have you signed your commits? Use git commit -s to sign.
  • Do all unittests finish successfully before sending PR?
    1. pytest or (if your machine does not have GPU) pytest --cpu from the root folder (given you marked your test cases accordingly @pytest.mark.run_only_on('CPU')).
    2. Sparrowhawk tests bash tools/text_processing_deployment/export_grammars.sh --MODE=test ...
  • If you are adding a new feature: Have you added test cases for both pytest and Sparrowhawk here.
  • Have you added __init__.py for every folder and subfolder, including data folder which has .TSV files?
  • Have you followed codeQL results and removed unused variables and imports (report is at the bottom of the PR in github review box) ?
  • Have you added the correct license header Copyright (c) 2023, NVIDIA CORPORATION & AFFILIATES. All rights reserved. to all newly added Python files?
  • If you copied nemo_text_processing/text_normalization/en/graph_utils.py your header's second line should be Copyright 2015 and onwards Google, Inc.. See an example here.
  • Remove import guards (try import: ... except: ...) if not already done.
  • If you added a new language or a new feature please update the NeMo documentation (lives in different repo).
  • Have you added your language support to tools/text_processing_deployment/pynini_export.py.

PR Type:

  • New Feature
  • Bugfix
  • Documentation
  • Test

If you haven't finished some of the above items you can still open "Draft" PR.

surendran-246 and others added 10 commits July 7, 2026 13:51
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
for more information, see https://pre-commit.ci

Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
@surendran-246 surendran-246 changed the title Tamil language cardinal semiotic class Tamil TN Cardinal Semiotic Class Jul 7, 2026
@surendran-246
surendran-246 marked this pull request as ready for review July 7, 2026 10:50
surendran-246 and others added 6 commits July 8, 2026 11:04
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
@github-actions

Copy link
Copy Markdown

This PR is stale because it has been open for 14 days with no activity. Remove stale label or comment or update or this will be closed in 7 days.

@github-actions github-actions Bot added the Stale label Jul 23, 2026
Signed-off-by: surendran-246 <surendrans@nvidia.com>
surendran-246 and others added 2 commits July 23, 2026 18:44
@github-actions github-actions Bot removed the Stale label Jul 24, 2026
Signed-off-by: surendran-246 <surendrans@nvidia.com>
@github-actions

github-actions Bot commented Aug 8, 2026

Copy link
Copy Markdown

This PR is stale because it has been open for 14 days with no activity. Remove stale label or comment or update or this will be closed in 7 days.

pre-commit-ci Bot and others added 2 commits August 19, 2026 07:23
Signed-off-by: surendran-246 <surendrans@nvidia.com>
@folivoramanh
folivoramanh requested a review from mgrafu August 28, 2026 03:45
@surendran-246

Copy link
Copy Markdown
Author

make sure both pytest and sparrrowhawk test pass

Completed the pytest and sparrowhawk tests, and both passed successfully.

@@ -0,0 +1,36 @@
100 நூறு

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

looking at one hundred, these could be compressed into a rule. would this also apply for the other hundreds? if so, please refactor

(1 > நூ) followed by (delete 00 or insert ற்) followed by (insert று)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, applies to all hundreds. Replaced hundred_ta.tsv with hundred_stem.tsv (digit → stem only).

Exact = stem + delete "00" + insert "று" (100, 200, … 900)
Combining = stem + insert "ற்" + insert "று" (101–999)
One rule now covers both forms from the same stem table.

teens_and_ties = teens_ties

# digit_oru
one_oru = pynini.cross("1", "ஒரு") | pynini.cross("௧", "ஒரு")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

let's have everything be files instead of hardcoding

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed and moved to one_oru.tsv file.

thousand_exact = thousand_stem + pynutil.insert("ம்")
thousand_prefix = thousand_stem + pynutil.insert("த்து")

self.digit = digit

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these don't seem to be used anywhere?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thousand_exact and thousand_prefix are used (they feed into graph_thousands), so they were kept. self.digit was genuinely unused, so it was removed.


single_digit = digit | zero
self.single_digits_graph = single_digit + pynini.closure(insert_space + single_digit)
zero_del = pynutil.add_weight(pynutil.delete(NEMO_ALL_ZERO), -0.1)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this weight absolutely necessary?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not necessary, and it has been removed.

self.graph_thousands = graph_thousands
tails = [single_digit, teens_ties, graph_hundreds, graph_thousands]

# TEN-THOUSANDS (10^4): stem + ஆயிரம்

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these should also be in tsv files. the process also repeats for graph_ten_thousands, graph_lakhs and graph_ten_lakhs so it could be compressed

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Compressed into a single add_scale() helper that wraps band() + tail-list append, so the same 3-line pattern isn't repeated for each scale level. Related hardcoded words (ஆயிரம்/ஆயிரத்து, லட்சம்/லட்சத்து) moved into scale.tsv.

graph_ten_lakhs, # ten-lakhs of crores
]
with open(get_abs_path("data/numbers/crore.tsv"), encoding="utf-8") as f:
crore_exact_raw, crore_prefix_raw = [line.strip() for line in f if line.strip()]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why not string_file?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

switched to pynini.string_file.


def __init__(self, deterministic: bool = True):
super().__init__(name="punctuation", kind="classify", deterministic=deterministic)
s = "!#%&\'()*+,-./:;<=>?@^_`{|}~\""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

does the language require any additional punctuation?

@surendran-246 surendran-246 Sep 3, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The tagger already covers all Unicode punctuation categories, and Tamil mostly reuses ASCII punctuation, so no additional punctuation is needed.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove this file

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove this file

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove this file

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed.

escaped_written=$(printf '%s' "$written" | sed 's/\\/\\\\/g')
denorm_pred=$(echo "$escaped_written" | normalizer_main --config=sparrowhawk_configuration.ascii_proto 2>&1 | tail -n 1 | sed 's/\xC2\xA0/ /g')

# trim white space

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove commented out code not needed

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed.

@@ -0,0 +1,53 @@
4 நான்குகள்~நான்கு நான்குகள்

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

does this cover all rule categories included? at the very least I see negative numbers missing

@surendran-246 surendran-246 Sep 3, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, it covers all rules, and I’ve added negative numbers and comma-grouped numerals to the test cases

class VerbalizeFinalFst(GraphFst):
"""
Finite state transducer that verbalizes an entire sentence, e.g.
tokens { name: "its" } tokens { time { hours: "twelve" minutes: "thirty" } } tokens { name: "now" } tokens { name: "." } -> its twelve thirty now .

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

all definitions should be in the language and follow the format of a real output

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed and updated the docstring example with real Tamil input and output.


self.optional_sign = pynini.cross("negative: \"true\"", "minus ")
if not deterministic:
self.optional_sign |= pynini.cross("negative: \"true\"", "negative ")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why is this in English?

@surendran-246 surendran-246 Sep 3, 2026

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — replaced the English negative with the Tamil equivalent எதிர்மறை.

def __init__(self, deterministic: bool = True):
super().__init__(name="cardinal", kind="verbalize", deterministic=deterministic)

self.optional_sign = pynini.cross("negative: \"true\"", "minus ")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why is this in English?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — replaced the English minus with the Tamil equivalent கழித்தல்.

self.optional_sign = pynini.cross("negative: \"true\"", "minus ")
if not deterministic:
self.optional_sign |= pynini.cross("negative: \"true\"", "negative ")
self.optional_sign |= pynini.cross("negative: \"true\"", "dash ")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why is this in English?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — replaced the English dash with the Tamil equivalent கோடுகுறி.

@mgrafu

mgrafu commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

please make sure that all new directories have init files

surendran-246 and others added 2 commits September 3, 2026 10:23
Signed-off-by: surendran-246 <surendrans@nvidia.com>
@surendran-246

Copy link
Copy Markdown
Author

Thanks for the review, @mgrafu. Addressed all the feedback comments.

Testing: Cardinal pytest passed; Sparrowhawk TNCardinal passed.

@surendran-246
surendran-246 requested a review from mgrafu September 3, 2026 06:08
surendran-246 and others added 5 commits September 3, 2026 17:33
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Signed-off-by: surendran-246 <surendrans@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants