My Favorite Bugs: Invalid Surrogate Pairs
george.mand.is
My Favorite Bugs: Invalid Surrogate Pairs
1–10 of 53 posts
Re: My Favorite Bugs: Invalid Surrogate Pairs
#2I recently ported a program from python to rust and the original author used string regexes. Input and output document encoding mattered but the characters that needed to be matched were always lower ASCII. The python program could have used binary regexes, but instead forced an input encoding (UTF-8) and made the user choose an output encoding. When the input comes from an unknown process or legacy data, however, you don’t always get the luxury of assuming the encoding. Switching to binary regexes and ignoring encoding altogether simplified logic, eliminated classes of errors, and made the program work in scenarios it couldn’t earlier. Getting rid of the last decoding/encoding code gave me so much relief, especially when all of the whacky encoding tests I had already written continued to work.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#3Author went for Intl.Segmenter too: https://github.com/cheeaun/phanpy/issues/1491
Re: My Favorite Bugs: Invalid Surrogate Pairs
#4- https://george.mand.is/invalid-surrogate-pairs/
I thought it was something that's easier to play with and feel than necessarily just read about.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#5Re: My Favorite Bugs: Invalid Surrogate Pairs
#6Was already bad enough that instead of bytes, we have to worry about code points. Now even that isn’t enough?
It would have been expensive, but all characters should have been fixed size 64bit values.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#7it's good to know about surrogate pairs in unicode. It was new to me too when being part of tracking down incomplete unicode flags in the (excellent) phanpy mastodon client. Author went for Intl.Segmenter too: https://github.com/cheeaun/phanpy/issues/1491
Re: My Favorite Bugs: Invalid Surrogate Pairs
#8Re: My Favorite Bugs: Invalid Surrogate Pairs
#9Great write-up. Do most modern languages handle invalid surrogates gracefully, or is it still a "good luck" situation depending on the runtime?
[0] But everyone disagrees as to what indexing a string means, so you need to make an actual choice if you want anything involving indexing to match across languages.
Re: My Favorite Bugs: Invalid Surrogate Pairs
#10Great write-up. Do most modern languages handle invalid surrogates gracefully, or is it still a "good luck" situation depending on the runtime?
It was really `encodeURIComponent` that didn't handle it gracefully.
If you just type this into the console (surrogate pair for cowboy smiley face emoji), you see it encodes it ("%F0%9F%A4%A0"):
encodeURIComponent("\uD83E\uDD20")
If you give it an invalid surrogate pair, it will throw an actual error:
encodeURIComponent("\uDD20\uD83E")