Live data from Hacker News

When Compilers Disagree About UTF‑8

nemanjatrifunovic.substack.com

1–8 of 8 posts

Re: When Compilers Disagree About UTF‑8

#2
You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop.

Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. This opens up even more state machine possibilities.

Re: When Compilers Disagree About UTF‑8

#3

You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop. Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII o…

> So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints.

Depends on what kind of text you're processing. Many languages use the ASCII range for spaces/newlines, digits, and punctuation.

Re: When Compilers Disagree About UTF‑8

#4
post #3

You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop. Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII o…

> So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII or emoji codepoints. Depends on what kind of text you're processing. Many languages use the ASCII range for spaces/newlines, digits, and punctuation.

That, plus markup, be it something XML/HTML-like or something Markdown-like.

Re: When Compilers Disagree About UTF‑8

#5

You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop. Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII o…

This is one of the optimizations from ClickHouse - detect ASCII and go a fast path: https://github.com/ClickHouse/ClickHouse/blob/6cde32de1a2463...

Re: When Compilers Disagree About UTF‑8

#6

You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop. Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII o…

FWIW I believe the second optimization is at odds with the first. It’s really hard to do this kind of conditional historical check in SIMD.

Re: When Compilers Disagree About UTF‑8

#8

You could actually use SIMD instructions to detect sequences of bytes with the top bit cleared, thus allowing bulk copies of ASCII text without a per-byte loop. Another optimization would be to take advantage of the fact that codepoint usage tends to cluster around the language of the text. So if you detect usage of 3-byte encodings, chances are you'll continue encountering only 3-byte encodings, with the odd ASCII o…

FWIW I believe the second optimization is at odds with the first. It’s really hard to do this kind of conditional historical check in SIMD.

[deleted]