ITADN

`\p{...}` gives wrong results when mixed with letters under `re.IGNORECASE`

#155297Openlkk7 创建于 14 天前
type-bugextension-modulestopic-regex
L
lkk7commented
### Bug description: `\p{Lu}` matches uppercase Unicode letters. It breaks when mixed with a standard letter with `re.IGNORECASE`: ```python import re property_only = re.compile(r"[\p{Lu}]", re.IGNORECASE) property_plus_one = re.compile(r"[\p{Lu}1]", re.IGNORECASE) property_plus_a = re.compile(r"[\p{Lu}a]", re.IGNORECASE) assert property_only.fullmatch("B") assert property_plus_one.fullmatch("B") assert property_plus_a.fullmatch("B"), "adding 'a' broke the match for 'B'" ``` Expected: adding `"a"` to the character class shouldn't break the existing match Actual: `AssertionError: adding 'a' broke the match for 'B'` ### CPython versions tested on: CPython main branch ### Operating systems tested on: macOS <!-- gh-linked-prs --> ### Linked PRs * gh-155299 <!-- /gh-linked-prs -->
1 条评论