Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Java ZipFile entries: reject duplicate names before selection

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A ZIP archive can contain more than one entry with the same name. A reader must define whether duplicates are rejected before choosing one entry by name.

Operational contract

The code reads central-directory entries without inflating them, caps the entry count at 47, and rejects duplicate names. It returns only a metadata list. This does not validate entry paths for extraction or prove the archive payloads are safe to inflate; those are separate steps. If a consumer uses ZipFile.getEntry after accepting duplicates, different tools may select or present a different occurrence. Rejecting ambiguity early keeps a manifest name tied to exactly one archive member under this policy.

Failure case

A supplier sends an archive with two entries both named receipt-47.xml, one approved and one altered. A UI might display one while a later reader opens the other. This scan fails the archive before either entry is parsed. It also fails a 48-entry archive rather than silently ignoring the tail.

Java code

Java
import java.io.IOException;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.Enumeration;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import java.util.zip.ZipEntry;
import java.util.zip.ZipFile;

public class UniqueZipEntryNames {
    public static List<String> scan(Path archive) throws IOException {
        List<String> names = new ArrayList<>();
        Set<String> observed = new HashSet<>();
        try (ZipFile zip = new ZipFile(archive.toFile())) {
            Enumeration<? extends ZipEntry> entries = zip.entries();
            while (entries.hasMoreElements()) {
                if (names.size() == 47) throw new IOException("ZIP entry count exceeds cap");
                String name = entries.nextElement().getName();
                if (!observed.add(name)) throw new IOException("Duplicate ZIP entry: " + name);
                names.add(name);
            }
        }
        return names;
    }
}

Performance and ownership cost

Scanning E metadata entries is O(E) expected time and O(E) memory for names and the set, capped at 47. It does not incur the cost of decompression. A later payload read needs its own expanded-byte bound.

Common Mistakes

  • Do not select by name before resolving duplicates.
  • Do not mistake this metadata scan for extraction-path validation.
  • Do not trust entry count or size metadata as a decompression limit.

Connected lessons

java
archive and compression boundaries
zip-duplicate-entry-names
Storage details