Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> It has nothing to do with UTF32

Right, my point is that the concept of a code point is seldom useful unless doing storage stuff with utf32 or implementing unicode algorithms. Python may expose an API of code points but that doesn't mean that it's meaningful.

Performance arguments can be made as to why the API should use code points instead of grapheme clusters, so there are legitimate reasons for Python (and Rust, and many other languages) to do so. Sometimes you just need some comparable notion of length and "number of code points" is acceptable.

However, you should be careful when writing code that confers meaning to the concept of a code point. A lot of code does this (using code points when they mean glyphs or grapheme clusters).



  > Right, my point is that the concept of a code point
  > is seldom useful unless doing storage stuff with 
  > utf32 or implementing unicode algorithms.
In XML land, where strings are almost always UTF8, XPath offers a string-to-codepoints() function that returns a sequence of integers, and a corresponding codepoints-to-string(). These two have been invaluable to me on many occasions when doing string manipulation gymnastics.


> when doing string manipulation gymnastics.

What kind of string manipulation gymnastics? I'd be wary of using codepoints for string manipulation for anything other than algorithms where you are explicitly asked to (e.g. algorithms that implement operations from the unicode spec)


On the other hand, there are algorithms embedded in widely-deployed standards which are defined in terms of code points.

For example, one I know quite well from having implemented it in Python: the HTML5 color parsing algorithm (the one that turns even incredible junk strings like "chucknorris" into color values) requires, in step 7 of the parsing process, replacing any code point higher than U+FFFF with the sequence '00' (that's two instances of U+0030 DIGIT ZERO).

And personally I think code points, as the basic atomic units of Unicode, do make sense as the things strings are made up of; I wish Python had better support for identifying graphemes without third-party libraries, but since Unicode encodings all map back to code points it makes sense to me that a Unicode string is a sequence of those rather than a sequence of some more-complex concept.


> there are algorithms embedded in widely-deployed standards which are defined in terms of code points.

From my original comment:

> You only care about code points when dealing with UTF32 strings or when implementing operations on unicode text.

These operations fall in the latter. It's still pretty niche. If an algorithm is defined explicitly in terms of code points this makes sense. Stuff starts falling apart when people assign meaning to code points and use it as a placeholder for other concepts like glyph or columns of grapheme cluster.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: