My Favorite Bugs: Invalid Surrogate Pairs
11–20 of 53 posts
Re: My Favorite Bugs: Invalid Surrogate Pairs
#12Once I ran into this it became hard to treat strings “normally” in any situation or, alternatively, I’d force hard encoding requirements in the domain. Regardless, handling grapheme clusters properly is hard and easy to get wrong. I recently ported a program from python to rust and the original author used string regexes. Input and output document encoding mattered but the characters that needed to be matched were al…
If I'm remembering correctly, we briefly explored a solution where we told Python "This is a UTF-16LE encoded string" so the count would match, but I think we learned/realized the endianness is actually dictated by the client's machine (Going from memory here). Ultimately we just changed the solution so the client was the source of truth about lengths and counts.
These threads are surfacing all kinds of things I forgot about and didn't add in that blog post. Maybe I need to write another, haha.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#13Damn, I’ve never really had to deal with Unicode all that much. Was already bad enough that instead of bytes, we have to worry about code points. Now even that isn’t enough? It would have been expensive, but all characters should have been fixed size 64bit values.
It would have been a non-starter, and then we'd all be dealing with Shift-JIS, BIG5, and FSM knows how many different codepages to this day. UTF-8 is about as elegant as it gets, though Java and JS still managed to fuck that up too (they both encode every codepoint outside the BMP as surrogate pairs in UTF-8)
Re: My Favorite Bugs: Invalid Surrogate Pairs
#14Damn, I’ve never really had to deal with Unicode all that much. Was already bad enough that instead of bytes, we have to worry about code points. Now even that isn’t enough? It would have been expensive, but all characters should have been fixed size 64bit values.
You're making the same mistake that numerous people made before you: thinking that it's as simple as using arrays of large enough numbers. First they thought that two bytes per symbol would be enough, then four. Spoiler alert: it wasn't. And eight won't work either.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#15emojies are a sequence of Unicode codepoints producing a single grapheme. Splitting in the middle of a grapheme will produce two valid strings, but with some funky half baked emoji. So for a text editor it makes sense to split between grapheme boundaries.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#16Great write-up. Do most modern languages handle invalid surrogates gracefully, or is it still a "good luck" situation depending on the runtime?
Modern string libraries largely use UTF-8 [0], and surrogates, regardless of whether they’re paired, are invalid in UTF-8. So, in a modern string library, as built in to most modern languages, you will not encounter surrogates except when translating between encodings. [0] But everyone disagrees as to what indexing a string means, so you need to make an actual choice if you want anything involving indexing to match a…
Java did not get the memo. Since the char type is fixed at 16 bits, it uses surrogates to encode everything outside the BMP, regardless of the encoding.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#17Damn, I’ve never really had to deal with Unicode all that much. Was already bad enough that instead of bytes, we have to worry about code points. Now even that isn’t enough? It would have been expensive, but all characters should have been fixed size 64bit values.
> It would have been expensive, but all characters should have been fixed size 64bit values You're making the same mistake that numerous people made before you: thinking that it's as simple as using arrays of large enough numbers. First they thought that two bytes per symbol would be enough, then four. Spoiler alert: it wasn't. And eight won't work either.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#18As for using extended grapheme clusters, it sounds a little bit iffy—maybe possible to use correctly, maybe not, because they’re not stable over time. That style of thing has created some fascinating bugs, like (a few years ago) index corruption in PostgreSQL due to collation changes.
Unicode scalar values are technically-safe: you can’t introduce invalid Unicode. But you can definitely still end up with nonsense.
> We made emoji an atomic node type.
That avoids problems for emoji, but leaves the underlying hazard untouched. I imagine it could still theoretically occur with other text, probably CJK. But probably only theoretically.
> This splits by grapheme clusters rather than code units. No orphaned surrogates, no split emoji. It's what .slice() should have been doing all along, but of course UTF-16 predates emoji by decades.
I do not agree that slice() should operate on extended grapheme clusters. Don’t lump the grapheme cluster/scalar value split in with the sins of UTF-16 and its unreliable code point/code unit split.
UTF-16 was unforced error (and I still can’t work out why it wasn’t obvious from the start that UCS-2 would never be enough). But the concept of multiple scalars contributing to the logical unit was always inevitable.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#19Damn, I’ve never really had to deal with Unicode all that much. Was already bad enough that instead of bytes, we have to worry about code points. Now even that isn’t enough? It would have been expensive, but all characters should have been fixed size 64bit values.
> It would have been expensive, but all characters should have been fixed size 64bit values. It would have been a non-starter, and then we'd all be dealing with Shift-JIS, BIG5, and FSM knows how many different codepages to this day. UTF-8 is about as elegant as it gets, though Java and JS still managed to fuck that up too (they both encode every codepoint outside the BMP as surrogate pairs in UTF-8)
Re: My Favorite Bugs: Invalid Surrogate Pairs
#20In summary, Unicode code points (characters) are 32 bit. JavaScript manipulates Unicode in utf-16 for historical reasons, because at some point before Unicode, 16 bit was deemed enough (ucs-2). utf-16 run length encodes Unicode 32 codepoints into one or two code units. Splitting in a middle of a codepoints produces one invalid half string, and one semantically different half string. emojies are a sequence of Unicode…
21-bit, actually. It was supposed to be 32-bit, but UTF-16 caps out at 21-bit, so they lopped eleven bits of potential from Unicode (and UTF-8, so no more six-byte encoding).
> at some point before Unicode
No, in the early days of Unicode.
> run length encodes
Um… what? RLE is a data compression thing, UTF-16 has nothing to do with it.