Normalization is very expensive and does not belong at the FS level. This kills performance for some classes of applications.
The comparison with other filesystems does not hold since applications for other OSs have always been developed with no normalization at FS level, and hence it was done by the applications, or through the use of high-level OS APIs. Mac applications, on the other hand, expect it to be the responsibility of the filesystem. This is explained pretty clearly in the article.
On a side note, I wish Apple would have taken this opportunity to switch normalization from NFD to NFC, which basically everything else uses. The distinction causes complexity and often issues in software which share data between Apple platforms and other platforms, such as version control systems for instance.
EDIT: according to pilif's comment, they did, which is awesome!
Do you have any recent benchmarks showing a significant impact from normalization? I haven't seen that on anything in at least a decade and that was simply Red Hat shipping an ancient and completely unoptimized libicu.
> Mac applications, on the other hand, expect it to be the responsibility of the filesystem. This is explained pretty clearly in the article.
Again, the article made a huge sweeping claim without supporting it. That's simply not true in either way – many apps on every OS don't handle this at all, some handle it consistently everywhere, and what a “Mac application” means varies widely from “clean, modern Cocoa” to “uses a lot of C, etc. libraries”, “cross platform C++”, “C# port”, “Electron shell”, etc. You can't make any statement which is true for every single one of those categories, much less for every code path which eventually results in a filesystem call. I've run into cases where something mostly worked until you hit their integrated ZIP, Git/SVN, etc. support and found a new way that a filename was constructed.
My point wasn't that everything is fine but simply that this is complicated and no decision results in avoiding problems. Not normalizing allows for confusing visually-identical files; normalizing results in errors or data loss which will be blamed on the OS.
> Rather than saying the article is wrong, can you demonstrate /why/ it is wrong?
I think you might to re-read my entire comment: note that I'm not arguing that the technical details are wrong, only that they're insufficient to support the huge “APFS is unusable” conclusion.
As previously noted, Windows and Linux work the same way and they are used by more people in individual non-English locales than the total number of Mac users. Would you say “NTFS is unusable by non-English users” is a useful statement?
There's plenty of room to say that a particular tool needs improvement, or that people making systems which copy or archive files should check for pathological cases, but it doesn't help anything to overstate the case so broadly.
The issue isn't a "bag-of-bytes" filename model. The issue is a "bag-of-bytes" filename model combined with an inconsistent normalization scheme.
It's not a problem on Windows or Linux filesystems because Windows and Linux don't provide a half-assed normalization scheme that lets me fairly easily create files that can't be accessed. If the Cocoa libraries did no normalization, then the resulting behavior might be obnoxious from a human-interface perspective, but I don't think the article would describe it as "little short of catastrophic".
I'm sitting here on my US English keyboard typing scancodes that look just like they did in 1990, so I'm not the best authority on how big of a problem it really is, but I'd guess it's going to result in a lot of bugs. Anyone who's ever tried to use a Mac with a case-sensitive HFS+ partition should be able to tell you that programmers can't even "normalize" their filenames consistently strictly within their native language.
> It's not a problem on Windows or Linux filesystems because Windows and Linux don't provide a half-assed normalization scheme that lets me fairly easily create files that can't be accessed
This is only true if you're talking about the kernel APIs. Unfortunately, filenames come from a variety of sources and it's easy to find tools which inconsistently normalize them – e.g. simply copying and pasting a name from a Word doc, web page, etc. which has different normalization than whatever originally created the file – or which produce either duplicate error messages or confusing error messages because the normalization form used in a file doesn't match the normalization form written on disk.
I've encountered variations of this problem on all three systems. No approach is going to handle 100% of the filenames in the wild and all of them will require extra care in the user-interface which may or may not have been done – e.g. the Windows Explorer still provides no way to tell why Café.txt and Café.txt are not the same file – and fixing the cases where programs are internally inconsistent. APFS switching will expose some programs which were unsafe before but since it's consistent with the other common filesystems it'll remove the need for every archive, version control, etc. system to either special-case or break.
Normalization is not that expensive if you only care to do normalization-insensitive string comparison and string hashing.
The reason it's not that expensive is: a) this requires no memory allocation, and b) most characters in most strings require no normalization!
Notionally you just look at pairs of next codepoints, and if the second one isn't combining and the first one is canonical for the chosen NF (a very fast check for ASCII!), then there's no need to normalize the first, otherwise you gather the combining codepoints and normalize, producing one normalized character and restarting the process where you left off. Most of the time the first codepoint requires no normalization, so the fast path is fast -- not as fast as a normal strcmp() or memcmp(), but still pretty fast.
In HFS+ it's even only done once per-create, so pretty cheap.
In ZFS it's once per-open()/stat()/and so on. But still, not at all on readdir(), and anyways, it's highly optimized. For an all ASCII filename the slow path is never taken, and for a mostly ASCII filename the slow path is only taken for non-ASCII codepoints that are followed by combining codepoints (that check is itself a slower-than-the-fast path, but still faster than the slowest path).
The comparison with other filesystems does not hold since applications for other OSs have always been developed with no normalization at FS level, and hence it was done by the applications, or through the use of high-level OS APIs. Mac applications, on the other hand, expect it to be the responsibility of the filesystem. This is explained pretty clearly in the article.
On a side note, I wish Apple would have taken this opportunity to switch normalization from NFD to NFC, which basically everything else uses. The distinction causes complexity and often issues in software which share data between Apple platforms and other platforms, such as version control systems for instance.
EDIT: according to pilif's comment, they did, which is awesome!