Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You can't do that with a .zip file because the file header information is actually at the end of the file. You could work around that by sending the header first, though.


You can't put the header first. It is an intentional feature of .zip that the header is at the end so you can update a large zip file by just appending a new header to the end. That way the entire file does not have to be rewritten. Just read the old header, append new files, append new header. This was important back in floppy disk days


>> Just read the old header, append new files, append new header. This was important back in floppy disk days

Don't forget 'overwrite old header'

Important for floppy disks in two ways, one because of space constraints and two, because of how slow floppies were.


Don't forget 'overwrite old header'

Why? The new one will become the 'real' one since it will be at the end of the modified file. So it doesn't really matter if you delete the old one. If you're really trying to squeeze file sizes down, you could reference the old 'header' from the new one, so that you do not have to list the entire archive's contents again.


Having written code to extract files from a ZIP file, it's because the header is variable sized, anywhere from 22 to 65,557 bytes in size (22 bytes fixed, up to 65,535 bytes for a comment). There are two ways to scan for the header, one is to seek just 22 bytes shy of the end of the file and start checking backwards (since the majority of ZIPs I've encountered do not have the comment) or, seek 65,557 bytes from the end and scan forward.

This fact has been used to construct pathological ZIP files where one tool will report one list of files and another tool will list a different set of files. That's why you really need to overwrite the old header.


then you wrote a bad extractor. Scanning forward is against the format. Zipping up a zip file is a perfectly valid thing to do. If the inner zip is stored uncompressed you'll have 2 headers at the end. If you scan forward you'll fail.


Why not? Isn't the purpose of a file compressor to squeeze file sizes down?


You don't have to overwrite the old header. Concat any zip files you want and feed them to conforming deconpressors and they correctly extract only the last headers files.

If you didn't only look at the last header then you'd have the issue that I can store fake headers as uncompressed content (type 0) and your unzip util would screw up


To clarify: if it's HTTP, you can request the end of file "header" first if the server supports the `Range` HTTP header (many servers support this when it's appropriate). So it's not doable by simple bash-style piping, but a dedicated tool can request the required bytes in the correct order.

I'm pretty sure I've used something similar, which was already implemented in the .NET framework. See their docs on constructing a ZipArchive object:

"If the underlying file or stream supports seeking, the files are read from the archive as they are requested. If the underlying file or stream does not support seeking, the entire archive is held in memory."

https://msdn.microsoft.com/en-us/library/system.io.compressi...


In fact .zip file contains file header information twice: before every compressed file (local file header) and at the end of the .zip file (central directory). It's possible to omit file size in local file header, but most most compressing utilities doesn't use this option. So most .zip files could be extracted without seeking to the end.


How do you send the header first? Is that a feature of the http framework?


It'd be something you'd have to write yourself, of course.

On the server side, read it and send it before the file. Maybe go through all your zips and copy the needed metadata to another location.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: