It's a good idea until you end up with two files that have the "same" name (eg. Amélie.jpg and Amélie.jpg) because one uses decomposed characters (U+0065 and U+0301) and the other one uses a single character (U+00E9).
If the difference is not visible in your browser (it shouldn't), try copy-pasting those two filenames in a text editor, one of them is 10 characters long and one of them is 11 characters long.
It can get worse - imagine a program that normalizes all paths before handing them off to the fs. Then you go to delete the denormalized file, but it deletes the normalized one. With this change, the only safe way to handle paths would be to NOT normalize at all, but Apple are recommending the exact opposite. I predict a lot of confusion will arise from this change.
The problem exists already on the visual level, since there are pairs of distinct Unicode characters that look very much alike.
I think the only thing you can really do about that is restricting to ASCII minus control characters. Probably many programmers would be willing to accept that, but non-technical users not.
The next best thing would be to enforce a canonical (code-point?) encoding. But given the complexities of Unicode that will get us only so far...
The software also doesn't consider the two strings given as an example above equivalent, and for the same exact reason: they're sequences of different code points.
>It's a good idea until you end up with two files that have the "same" name (eg. Amélie.jpg and Amélie.jpg) because one uses decomposed characters (U+0065 and U+0301) and the other one uses a single character (U+00E9).
It's still good then. And several systems allow for that just fine, including VMS (uniqueness comes from more than the filename). This is more alike real world folders (which can have identical items, e.g. two copies of the same paper), and is also an excellent way to keep different document versions (keep the name the same, change the date shown).
Are you also against case sensitive file systems? Otherwise you can end up with two files that have the "same" name - eg. anne.jpg and Anne.jpg). Does normalization cover such a case?
I think, the problem with Unicode normalization is, input methods are (well, conceptually, at least) meant to produce text, not binary. If filesystems are using binary data for filenames, there can be a case when is really no way to address a file by typing its name, even if you can type in that language. This isn't an issue for case sensitivity or alike.
>If filesystems are using binary data for filenames, there can be a case when is really no way to address a file by typing its name, even if you can type in that language
You could do that trivially in UNIX since forever.
Don't know about other UNIXes, but at least on GNU/Linux, neither IMEs nor filesystems are working with text data - it's all binary strings (with a few restrictions, like unacceptability of NULs). The only place where those binary strings are converted to text is when they're rendered (and this may cause some encoding-related oddities). So, sure one can do that.
I would've actually preferred for identifiers to be Unicode text strings.
Filesystems themselves treating file names and paths at bits of binary data is the right way to do it, normalization shouldn't be handled directly by the filesystem as it adds complexity and duplication of effort.
Ideally, the VFS layer should be handling all this garbage.
Programs often allow you to type arbitrary binary values using a keyboard, either by holding a keyboard control key and typing the numeric value, or prefixing it with '\'.
For ascii, I would gladly do this There may be edge cases outside of the English language that I'm not aware of.
I am somewhat willing to accept the value of case sensitivity in identifiers for programming languages (though it's often related to the verbosity of the language, as in Java: "Foo foo = new Foo()"), but not in file naming.
Of course, I realize that the ship has sailed, and we're stuck with case sensitivity in file systems.
If the difference is not visible in your browser (it shouldn't), try copy-pasting those two filenames in a text editor, one of them is 10 characters long and one of them is 11 characters long.