ITADN

Improve A-Z replace_tag definitions

#19Pull Requestfkoyer 创建于 2025-10-08
F
fkoyercommented
Problems with old definitions: * Tries to match UTF-8 and Latin-1 characters in same expression. e.g. \<A\> includes the byte sequence for "ã" in Latin-1 (\xE3) and UTF-8 (\xC3\xA3). This seems like a good thing at first but it can cause false positives if the text is in UTF-8 and the pattern is looking for Latin-1 * Contains redundant characters. e.g. \xE3 appears multiple times in \<A\> * Contains unnecessary characters. e.g. \xE3 also appears in \<V\> and \<Y\> * Patterns are case-insensitive. e.g. \<I\> attempts to match lowercase L but because it's case-insensitive, it also matches uppercase L * Some look-alike characters aren't matched e.g. \xEA\x93\xAE = LISU LETTER A (U+A4EE) Changes: * All byte sequences are UTF-8 only (no Latin-1) * All patterns are case-sensitive * Removed redundant and unnecessary characters * Added additional look-alike characters
合并状态:未合并 1 条评论