CharsetDecoder can report malformed and unmappable input as errors rather than silently substituting replacement characters.
Java CharsetDecoder: reject malformed UTF-8 instead of replacing bytes
Make the data-loss policy explicit
A replacement character can hide a damaged identifier or signature. Set both error actions to REPORT when the text is a protocol field that must round-trip exactly. If replacement is acceptable for display-only content, make that a separate path and label it accordingly.
The fixture decodes one valid identifier and rejects a malformed two-byte sequence. A strict UTF-8 reader covers file input; this lesson shows the decoder policy directly for in-memory bytes.
Keep character and byte lengths separate
A byte cap does not imply the same character cap. UTF-8 characters can occupy multiple bytes, and a Java String uses UTF-16 code units. Enforce each limit at the layer where it matters, before allocating or persisting.
Working program
import java.nio.ByteBuffer;
import java.nio.charset.CharacterCodingException;
import java.nio.charset.CodingErrorAction;
import java.nio.charset.StandardCharsets;
public class SupplierIdentifierDecoder {
static String decode(byte[] bytes) throws CharacterCodingException {
return StandardCharsets.UTF_8.newDecoder()
.onMalformedInput(CodingErrorAction.REPORT)
.onUnmappableCharacter(CodingErrorAction.REPORT)
.decode(ByteBuffer.wrap(bytes)).toString();
}
public static void main(String[] args) throws CharacterCodingException {
System.out.println(decode(new byte[] {'S', 'U', 'P', '-', '4', '7'}));
try { decode(new byte[] {(byte) 0xC3, (byte) 0x28}); }
catch (CharacterCodingException invalid) { System.out.println("invalid UTF-8 rejected"); }
}
}Output
SUP-47
invalid UTF-8 rejectedCost and ownership
Decoding is O(n) in input bytes and materializes characters in memory. A per-call decoder keeps mutable decoding state local; for a streaming pipeline, reuse one decoder within a single stream and handle underflow and overflow explicitly.
Common Mistakes
- Do not accept replacement characters silently in identifiers.
- Do not count UTF-8 bytes as Java characters.
- Do not share one mutable decoder across unrelated concurrent operations.
Read next
Java UTF-8 decoding: reject malformed bytes before parsing, bytebuffer flip compact state, Java file I/O: UTF-8, streaming reads, and path ownership, Java DecimalFormat: reject trailing input and pin the locale.
Continue with: Java DOM parsing: deny external XML access at the factory.
