Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"Unusable" is a strong word to use. Should filesystems be making up for our Unicode shortcomings? From a SW design perspective, is that the most sensible place to pass the burden of responsibility? I would say that another way to handle it is to store a file name as an array of bytes and put the burden on software developers to interpret Unicode correctly. Swift does this pretty nicely.

I would say the only downside to this approach, is a user wouldn't be able to distinguish two files with the same name apart, but it's hard to imagine how they'd get to creating such a situation in the first place without the developer rule above being violated.



> Should filesystems be making up for our Unicode shortcomings?

Absolutely, yes. File names are text by their very definition; that we've been treating them as "bags of bytes" is a historical tragedy. At the very least, file names need to be displayed, as text, to the user, so they should be stored as text, that is in some well-defined encoding, and yes, it should be the job of the filesystem driver / kernel to enforce that it's not writing garbage out to disk.

Reinventing that wheel in every system that in any way interacts with the filesystem is bad engineering, and doomed to fail.

Further, I don't see why the typical user should need to know or understand the differences between 'e\N{COMBINING ACUTE ACCENT}' and '\N{LATIN LOWERCASE E WITH ACUTE ACCENT}'. Likewise, I don't see why each and every piece of code should be forced to handle that. Developers will get this wrong. In fact, the article seems to say even Apple can't get it right, in that Finder will not correctly show the directory contents in some instances, and fails to open files in some instances, telling users the file "doesn't exist".


But what is text? Not everyone wants to use unicode. It is dependent of the platform, the region, the OS and on many other different things like LC_* variables on linux. Why should a filesystem depend on those too?


> and on many other different things like LC_ variables on linux.*

The point is that it doesn't need to be. I would entertain that not everyone might not want to use Unicode: in that case, the FS should still have a well-defined encoding, such that I can still arrive at a string to display to the user. The point is not that "Unicode is best" but that storing file names as "bags of bytes" is incredibly user unfriendly, and there needs to be a straightforward, no bullshit method to display and transmit the names of files.

But I would also argue that Unicode is the best we've got presently, and it would be pragmatic for a filesystem to simply adopt it outright. It's overwhelmingly the dominant character set in use today, especially if you ignore deprecated junk that Unicode is a strict superset of.


ZFS has a per-dataset option to allow/reject non-UTF-8. For valid UTF-8 names ZFS implements normalization-preserving/insensitive behavior. That was the best compromise we could find, and it works really well.


RDBMS's have gone through the same journey. First it was hardcoded, now for many of them we can specify the codepage or UTF encoding and collation of a DB, some even offer it per table or column, and so on.

Filesystems should do the same. Either pick one way and stick to it (expose it as UTF8, with a sensible normalization and collation) or offer several options that can be specified when creating a volume and do on the fly conversions for clients that need it.


Unicode is meant to replace all the other codesets. It's time to move on from them, except for compatibility reasons.

Anyways, it's possible to allow UTF-8 and non-UTF-8 on the same filesystem, and still provide form-preserving/insensitive behavior... ZFS does it.


> Should filesystems be making up for our Unicode shortcomings?

In this case: yes. Specifically the FS should implement normalization-preserving/insensitive behavior. http://cryptonector.com/2010/04/on-unicode-normalization-or-...




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: