Each successful Matcher.find advances the matcher search state, including when a pattern can match an empty span.
Java Matcher.find: account for the scan cursor and zero-length matches
Read positions as spans
An insertion-point pattern can match between characters and at the end. The two numbers from start and end are UTF-16 offsets delimiting a half-open span; they are not byte positions. A zero-width match has equal offsets.
The program prints the positions for an input of two ASCII characters. Java advances its search after an empty match, so the loop terminates. Unicode encoding explains why non-ASCII text can make a UTF-16 offset differ from a user-visible character count.
Reset for another pass
Once the search is exhausted, a subsequent find call does not start from the beginning. Call reset or create a fresh matcher when another full pass is required. Do not hand-roll an extra cursor increment on the same matcher without testing endpoint behavior.
Working program
import java.util.regex.Matcher;
import java.util.regex.Pattern;
public class InsertionPointScanner {
public static void main(String[] args) {
Matcher boundaries = Pattern.compile("(?=.)|$").matcher("AB");
while (boundaries.find()) {
System.out.println(boundaries.start() + ":" + boundaries.end());
}
}
}Output
0:0
1:1
2:2Cost and ownership
This fixture visits three positions with constant matcher storage. In general, cost depends on the pattern and input; zero-width matches can multiply output records, so cap emitted matches for untrusted text.
Common Mistakes
- Do not treat start and end as byte offsets.
- Do not call group before a successful find.
- Do not assume an empty match leaves find stuck forever.
Read next
Java Matcher.matches and find: whole-input validation versus extraction, regex group state, Java Unicode: code units, code points and UTF-8 bytes, Java strings and content equality.
