This is good technical content, but it's obvious that an AI wrote it.
Don't stop early: Case-folding source code at memory speed
31–40 of 40 posts
Re: Don't stop early: Case-folding source code at memory speed
#32TLDR: they implemented case folding with a lot more SIMD via autovectorization. > almost every fold preserves the UTF-8 length or shrinks it, but two outliers grow—U+023A (Ⱥ) and U+023E (Ɀ) are 2 bytes each yet fold to 3-byte characters (ⱥ, ɀ) Fix this by reversing it. Fold ⱥ to Ⱥ instead of the other way around. The search index won't only consist of lowercase characters any more, but that never mattered.
Re: Don't stop early: Case-folding source code at memory speed
#33Earlier quoted context omitted.
Yes. > This is genuinely interesting Are you sure you're not an LLM yourself?
Has this become a ‘smell’? I tend to start my HN comments that are going be negative with versions of this, or my habitual “Genuine question, ..”; but if this is going to flag readers’ internal LLM-detector I will have to find some other way to indicate I’m actually interested in a dialogue (versus the shit posting that a more brief reply might signal). I was talking with a junior at the office today about LLM output…
Re: Don't stop early: Case-folding source code at memory speed
#34TLDR: they implemented case folding with a lot more SIMD via autovectorization. > almost every fold preserves the UTF-8 length or shrinks it, but two outliers grow—U+023A (Ⱥ) and U+023E (Ɀ) are 2 bytes each yet fold to 3-byte characters (ⱥ, ɀ) Fix this by reversing it. Fold ⱥ to Ⱥ instead of the other way around. The search index won't only consist of lowercase characters any more, but that never mattered.
It does not look like the author posted this here since it was posted on the blog last week. So would probably be a better idea to suggest / recommend this at https://github.com/github/rust-gems/tree/main/crates/casefol...
Re: Don't stop early: Case-folding source code at memory speed
#35There’s some interesting information in there. Unfortunately the person or LLM writing this got pretty confused right in the introduction already. > “Suppose […] they type straße and you’ve stored STRASSE. To make these count as matches, you need […]” Really bad example, because as the article says later on, this casefold crate won’t match those two strings because the ß → ss conversion isn’t done. > “[str::to_lowerc…
Or if you really want to do it, your case-folding needs to map certain lowercase characters to to other lowercase characters (lowercase ß to ss, lowercase ı to i, etc), losing some of the meaning
Re: Don't stop early: Case-folding source code at memory speed
#36There’s some interesting information in there. Unfortunately the person or LLM writing this got pretty confused right in the introduction already. > “Suppose […] they type straße and you’ve stored STRASSE. To make these count as matches, you need […]” Really bad example, because as the article says later on, this casefold crate won’t match those two strings because the ß → ss conversion isn’t done. > “[str::to_lowerc…
If anything, the word STRASSE is a great example why case-folding followed by string comparison is not the same as a case-independent string comparison Or if you really want to do it, your case-folding needs to map certain lowercase characters to to other lowercase characters (lowercase ß to ss, lowercase ı to i, etc), losing some of the meaning
Re: Don't stop early: Case-folding source code at memory speed
#37There’s some interesting information in there. Unfortunately the person or LLM writing this got pretty confused right in the introduction already. > “Suppose […] they type straße and you’ve stored STRASSE. To make these count as matches, you need […]” Really bad example, because as the article says later on, this casefold crate won’t match those two strings because the ß → ss conversion isn’t done. > “[str::to_lowerc…
If anything, the word STRASSE is a great example why case-folding followed by string comparison is not the same as a case-independent string comparison Or if you really want to do it, your case-folding needs to map certain lowercase characters to to other lowercase characters (lowercase ß to ss, lowercase ı to i, etc), losing some of the meaning
s/case-folding/lowercasing/
Proper Unicode case-folding absolutely does map ß to ss, ς to σ, etc. Moreover, there are some scripts (IIRC Georgian) where for historical reasons case-folding yields uppercase letters, not lowercase ones. The case-folding mapping is specifically designed in concert with the comparison rules to yield the same result, that’s why it’s a separate operation from lowercasing.
(I believe the thing described in TFA is supposed to be proper case-folding in that sense, but given TFA is AI-written I wouldn’t trust its descriptions either way.)
That said, if you want to match the sort order customary in a specific language, you need to use language-specific rules for producing collation keys rather than generic case-folding. There’s no way out of this because different users of e.g. the Latin alphabet want contradictory results. And if you think you do want generic casefolding, then you probably actually want NFKC_Casefold instead unless your input is pre-normalized.
Re: Don't stop early: Case-folding source code at memory speed
#38Earlier quoted context omitted.
Yes. > This is genuinely interesting Are you sure you're not an LLM yourself?
I thought I wasn't, but you're making me second-guess myself.
Re: Don't stop early: Case-folding source code at memory speed
#39Re: Don't stop early: Case-folding source code at memory speed
#40This is good technical content, but it's obvious that an AI wrote it.
Agreed. This is genuinely interesting content, but there is no doubt in my mind that "The two operations diverge on real characters—ß, İ, final sigma—which is why lowercasing as a stand-in silently produces wrong matches." is LLM output. Are we doomed to spend the rest of our professional and personal lives reading AI output?