A ZIP archive can contain more than one entry with the same name. A reader must define whether duplicates are rejected before choosing one entry by name.
Java ZipFile entries: reject duplicate names before selection
Operational contract
The code reads central-directory entries without inflating them, caps the entry count at 47, and rejects duplicate names. It returns only a metadata list. This does not validate entry paths for extraction or prove the archive payloads are safe to inflate; those are separate steps. If a consumer uses ZipFile.getEntry after accepting duplicates, different tools may select or present a different occurrence. Rejecting ambiguity early keeps a manifest name tied to exactly one archive member under this policy.
Failure case
A supplier sends an archive with two entries both named receipt-47.xml, one approved and one altered. A UI might display one while a later reader opens the other. This scan fails the archive before either entry is parsed. It also fails a 48-entry archive rather than silently ignoring the tail.
Java code
import java.io.IOException;
import java.nio.file.Path;
import java.util.ArrayList;
import java.util.Enumeration;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import java.util.zip.ZipEntry;
import java.util.zip.ZipFile;
public class UniqueZipEntryNames {
public static List<String> scan(Path archive) throws IOException {
List<String> names = new ArrayList<>();
Set<String> observed = new HashSet<>();
try (ZipFile zip = new ZipFile(archive.toFile())) {
Enumeration<? extends ZipEntry> entries = zip.entries();
while (entries.hasMoreElements()) {
if (names.size() == 47) throw new IOException("ZIP entry count exceeds cap");
String name = entries.nextElement().getName();
if (!observed.add(name)) throw new IOException("Duplicate ZIP entry: " + name);
names.add(name);
}
}
return names;
}
}Performance and ownership cost
Scanning E metadata entries is O(E) expected time and O(E) memory for names and the set, capped at 47. It does not incur the cost of decompression. A later payload read needs its own expanded-byte bound.
Common Mistakes
- Do not select by name before resolving duplicates.
- Do not mistake this metadata scan for extraction-path validation.
- Do not trust entry count or size metadata as a decompression limit.
Connected lessons
- Java ZIP inputs: bounded staging and rejected path traversal
- Java ZIP entry extraction: cap expanded bytes while reading
- Java ZIP CRC32: detect corruption without claiming authenticity
- Java GZIPInputStream: cap bytes after decompression
- Java GZIPOutputStream: close before using the compressed bytes
- Java ZipOutputStream: close each entry and the archive before publication
- Java XML and archive boundaries quiz
- Advanced Java
