Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Java GZIPInputStream: cap bytes after decompression

Last updated: 4 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

GZIPInputStream yields uncompressed bytes. A compressed input-size cap does not bound the number of bytes a decoder can produce.

Operational contract

The method caps compressed input at 47,000 bytes and refuses more than 47,000 inflated bytes while reading chunks. It closes the decompressor even when a read or size check fails. It returns only a complete, accepted byte array. The inflated-byte limit protects application memory but not all CPU work consumed before reaching it; admission, time, and nesting policies may also be needed for untrusted archives. A checksum inside GZIP detects accidental corruption, not sender identity.

Failure case

A depot receives a 900-byte compressed manifest that expands to 82,000 bytes. Checking only the upload length would accept it. This decoder stops when the inflated stream crosses 47,000 bytes and does not pass a partial manifest to the parser.

Java code

Java
import java.io.ByteArrayInputStream;
import java.io.ByteArrayOutputStream;
import java.io.IOException;
import java.util.zip.GZIPInputStream;

public class BoundedManifestInflater {
    public static byte[] inflate(byte[] compressed) throws IOException {
        if (compressed.length > 47_000) throw new IOException("Compressed input exceeds cap");
        ByteArrayOutputStream result = new ByteArrayOutputStream();
        try (GZIPInputStream input = new GZIPInputStream(new ByteArrayInputStream(compressed))) {
            byte[] chunk = new byte[8_192];
            int count;
            while ((count = input.read(chunk)) != -1) {
                if (count > 47_000 - result.size()) throw new IOException("Inflated data exceeds cap");
                result.write(chunk, 0, count);
            }
        }
        return result.toByteArray();
    }
}

Performance and ownership cost

Decompression is O(I+O) for compressed input I and inflated output O in ordinary workloads. Retained output is O(O), capped here, with an 8,192-byte work buffer. Highly compressible input can consume disproportionate work before the cap is hit.

Common Mistakes

  • Do not use compressed length as an expanded-size limit.
  • Do not return a prefix after the limit is crossed.
  • Do not treat the GZIP checksum as authentication.

Connected lessons

java
archive and compression boundaries
gzip-inflated-byte-cap
Storage details