Java regex predefined word classes use an ASCII-oriented default in Java 8; UNICODE_CHARACTER_CLASS changes the predefined class behavior for Unicode text.
Java regex Unicode classes: decide what a word character means
Declare the character policy
A customer identifier containing an accented letter is not accepted by an ASCII-only word rule. The same regex source can accept it when compiled with UNICODE_CHARACTER_CLASS. Neither setting is universally correct: protocol identifiers may require ASCII, while human names need a more deliberate policy than word characters alone.
The fixture compares one identifier under both flags. Text normalization handles canonically equivalent spellings; changing regex flags does not normalize the input. Unicode representation covers code points and encoded bytes.
Avoid vague name validation
Names contain spaces, punctuation, scripts and conventions that a word-character rule misses. Keep the example as an identifier policy. If stored identifiers must compare consistently, define normalization and case rules separately before persistence.
Working program
import java.util.regex.Pattern;
public class CustomerCodeCharacterPolicy {
public static void main(String[] args) {
String customerCode = "Café_47";
Pattern ascii = Pattern.compile("\\w+");
Pattern unicode = Pattern.compile("\\w+", Pattern.UNICODE_CHARACTER_CLASS);
System.out.println("ascii=" + ascii.matcher(customerCode).matches());
System.out.println("unicode=" + unicode.matcher(customerCode).matches());
}
}Output
ascii=false
unicode=trueCost and ownership
Both patterns scan this short input with small state. More permissive character classes can admit values that downstream systems reject; validate the full storage and protocol contract rather than assuming a Unicode flag settles it.
Common Mistakes
- Do not assume the default word class covers every letter.
- Do not treat a Unicode class flag as text normalization.
- Do not use a generic word rule as a complete human-name validator.
Read next
Java text normalization: equal labels and explicit comparison policy, Java Unicode: code units, code points and UTF-8 bytes, Java Matcher.matches and find: whole-input validation versus extraction, Java strings and content equality.
