Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Confusable characters look similar or the same to humans.

Canonically equivalent Unicode sequences look the same to machines.

The latter is a much more significant problem, because it can wreak havoc with interoperability.



Very true - and it's amplified by inconsistency allowing problems to spread further before being noticed. At work I deal with a lot of Bag-It archives where we have a text manifest of checksums accompanying files on disk and this reliably bites users of tools when something (archive or network transfer tool, Git or SVN, transition to/from a Mac with HFS+, etc.) causes the encoding in the manifest not to match the local filesystem, and the confusion is amplified because some tools will handle normalization differences so the bug report is “why does tool A say this file is missing when Explorer/Finder and tool B say it's fine?”


Canonically equivalent Unicode sequences look the same to machines.

Memcmp disagrees, as do the default equality operators of most programming languages in existence.


Sure, but normalisation can nonetheless happen automatically and implicitly in many places.


Rust uses separate string type for file names. I think, that's a good approach. If language normalizes strings behind your back, that's not very good.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: