unicode/norm: do not compose across a ccc=0 character that combines backward - #69
unicode/norm: do not compose across a ccc=0 character that combines backward#69tannevaled wants to merge 1 commit into
Conversation
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
7691338 to
eab86a8
Compare
|
This PR (HEAD: eab86a8) has been imported to Gerrit for code review. Please visit Gerrit at https://go-review.googlesource.com/c/text/+/821020. Important tips:
|
|
Message from Gopher Robot: Patch Set 1: (1 comment) Please don’t reply on this GitHub thread. Visit golang.org/cl/821020. |
|
Message from Gopher Robot: Patch Set 1: Congratulations on opening your first change. Thank you for your contribution! Next steps: Most changes in the Go project go through a few rounds of revision. This can be During May-July and Nov-Jan the Go project is in a code freeze, during which Please don’t reply on this GitHub thread. Visit golang.org/cl/821020. |
…ackward UAX golang#15, D115: a character C is blocked from starter A when some B between them has ccc(B) = 0 or ccc(B) >= ccc(C). reorderBuffer.compose tracks the current starter in s, but only ever updates it inside the combinesBackward branch. That is enough for a ccc = 0 character that does not combine backward, because such a character ends the segment and never shares a buffer with an earlier starter. One that does combine backward has to stay in the buffer, since it may compose with what precedes it; when it does not, it falls through with s still pointing at the earlier starter, and the next mark composes across it. norm.NFC.String("iা̖̀") got "ìা̖" -- i + U+0300 composed across U+09BE want "iা̖̀" Set s when the character just written is itself a starter. The bug is not new and not tied to a particular Unicode version. Of the 35 characters in Unicode 17 that have ccc = 0 and are the second element of a canonical composition, NFC(<i, S, U+0300, U+0316>) fails to keep the "i" for 24 of them under the Unicode 15 tables and 33 under Unicode 17 -- among them U+09BE Bengali, U+0BBE Tamil, U+0D3E Malayalam, U+1B35 Balinese and U+102E Myanmar. The nine that differ between the two table sets are simply those Unicode 15 had not assigned yet. The two that never fail, U+0FB5 and U+0FB7, are composition exclusions. The added test uses U+09BE, which is assigned with ccc = 0 in both table sets, so it does not depend on which toolchain selects which tables, and it is paired with U+0903 -- ccc = 0, does not combine backward -- so the two cases sit either side of the distinction. Verified beyond the package's own tests: NFC, NFKC and NFD of 3112063 inputs were compared before and after, and every one of the 2778 (Unicode 15) and 4153 (Unicode 17) differing results was judged by a separate UAX golang#15 implementation built from UnicodeData.txt, itself checked against NormalizationTest.txt at 305184 and 320544 assertions with no failures. All of the differences are corrections; there are no regressions. Fixes golang/go#81001
eab86a8 to
aa7e629
Compare
|
This PR (HEAD: aa7e629) has been imported to Gerrit for code review. Please visit Gerrit at https://go-review.googlesource.com/c/text/+/821020. Important tips:
|
|
Message from t hepudds: Patch Set 2: (1 comment) Please don’t reply on this GitHub thread. Visit golang.org/cl/821020. |
UAX #15, D115: a character C is blocked from starter A when some B between
them has ccc(B) = 0 or ccc(B) >= ccc(C). reorderBuffer.compose tracks the
current starter in s, but only ever updates it inside the combinesBackward
branch. That is enough for a ccc = 0 character that does not combine
backward, because such a character ends the segment and never shares a
buffer with an earlier starter. One that does combine backward has to stay
in the buffer, since it may compose with what precedes it; when it does not,
it falls through with s still pointing at the earlier starter, and the next
mark composes across it.
Set s when the character just written is itself a starter.
The bug is not new and not tied to a particular Unicode version. Of the 35
characters in Unicode 17 that have ccc = 0 and are the second element of a
canonical composition, NFC(<i, S, U+0300, U+0316>) fails to keep the "i" for
24 of them under the Unicode 15 tables and 33 under Unicode 17 -- among them
U+09BE Bengali, U+0BBE Tamil, U+0D3E Malayalam, U+1B35 Balinese and U+102E
Myanmar. The nine that differ between the two table sets are simply those
Unicode 15 had not assigned yet. The two that never fail, U+0FB5 and U+0FB7,
are composition exclusions.
The added test uses U+09BE, which is assigned with ccc = 0 in both table
sets, so it does not depend on which toolchain selects which tables, and it
is paired with U+0903 -- ccc = 0, does not combine backward -- so the two
cases sit either side of the distinction.
Verified beyond the package's own tests: NFC, NFKC and NFD of 3112063 inputs
were compared before and after, and every one of the 2778 (Unicode 15) and
4153 (Unicode 17) differing results was judged by a separate UAX #15
implementation built from UnicodeData.txt, itself checked against
NormalizationTest.txt at 305184 and 320544 assertions with no failures. All
of the differences are corrections; there are no regressions.
Fixes golang/go#81001