Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Java CharsetDecoder: reject malformed UTF-8 instead of replacing bytes

Last updated: 5 Oct 20264 min read
tutorial
IntermediateBy AITrove Editorial

CharsetDecoder can report malformed and unmappable input as errors rather than silently substituting replacement characters.

Make the data-loss policy explicit

A replacement character can hide a damaged identifier or signature. Set both error actions to REPORT when the text is a protocol field that must round-trip exactly. If replacement is acceptable for display-only content, make that a separate path and label it accordingly.

The fixture decodes one valid identifier and rejects a malformed two-byte sequence. A strict UTF-8 reader covers file input; this lesson shows the decoder policy directly for in-memory bytes.

Keep character and byte lengths separate

A byte cap does not imply the same character cap. UTF-8 characters can occupy multiple bytes, and a Java String uses UTF-16 code units. Enforce each limit at the layer where it matters, before allocating or persisting.

Working program

Java
import java.nio.ByteBuffer;
import java.nio.charset.CharacterCodingException;
import java.nio.charset.CodingErrorAction;
import java.nio.charset.StandardCharsets;

public class SupplierIdentifierDecoder {
    static String decode(byte[] bytes) throws CharacterCodingException {
        return StandardCharsets.UTF_8.newDecoder()
                .onMalformedInput(CodingErrorAction.REPORT)
                .onUnmappableCharacter(CodingErrorAction.REPORT)
                .decode(ByteBuffer.wrap(bytes)).toString();
    }

    public static void main(String[] args) throws CharacterCodingException {
        System.out.println(decode(new byte[] {'S', 'U', 'P', '-', '4', '7'}));
        try { decode(new byte[] {(byte) 0xC3, (byte) 0x28}); }
        catch (CharacterCodingException invalid) { System.out.println("invalid UTF-8 rejected"); }
    }
}

Output

Output
SUP-47
invalid UTF-8 rejected

Cost and ownership

Decoding is O(n) in input bytes and materializes characters in memory. A per-call decoder keeps mutable decoding state local; for a streaming pipeline, reuse one decoder within a single stream and handle underflow and overflow explicitly.

Common Mistakes

  • Do not accept replacement characters silently in identifiers.
  • Do not count UTF-8 bytes as Java characters.
  • Do not share one mutable decoder across unrelated concurrent operations.

Read next

Java UTF-8 decoding: reject malformed bytes before parsing, bytebuffer flip compact state, Java file I/O: UTF-8, streaming reads, and path ownership, Java DecimalFormat: reject trailing input and pin the locale.

Continue with: Java DOM parsing: deny external XML access at the factory.

java
io
charsetdecoder-report-policy
Storage details