Live data from Hacker News

Charset="WTF-8"

wtf-8.xn--stpie-k0a81a.com

61–70 of 463 posts

Re: Charset="WTF-8"

#61
post #53

Earlier quoted context omitted.

If you just use the {Alphabetic} Unicode character class (100K code points), together with a space, hyphen, and maybe comma, that might get you close. It includes diacritics. I'm curious if anyone can think of any other non-alphabetic characters used in legal names around the world, in other scripts? I wondered about numbers, but the most famous example of that has been overturned: "Originally named X Æ A-12, the chi…

There’s this individual’s name which involves a clock sound: Nǃxau ǂToma[1] [1] https://en.m.wikipedia.org/wiki/N%25C7%2583xau_%C7%82Toma

Click characters are part of {Alphabetic}!

https://en.wikipedia.org/wiki/Click_consonant

https://www.compart.com/en/unicode/category/Lo

https://stackoverflow.com/a/4843363

Re: Charset="WTF-8"

#62
post #44
post #41

Earlier quoted context omitted.

I wonder how many of those packages end up in Vada, Italy. Or Cody, Wyoming. Or Buda, Texas...

I imagine the “Poland” part of the address would narrow it down somewhat.

I got curious if I can get data to answer that, and it seems so.

Based on xlsx from [0], we got the following ??d? localities in Poland:

1 x Bądy, 1 x Brda, 5 x Buda, 120 x Budy, 4 x Dudy, 1 x Dydy, 1 x Gady, 1 x Judy, 1 x Kady, 1 x Kadź, 1 x Łada, 1 x Lady, 4 x Lądy, 2 x Łady, 1 x Lęda, 1 x Lody, 4 x Łódź, 1 x Nida, 1 x Reda, 1 x Redy, 1 x Redz, 74 x Ruda, 8 x Rudy, 12 x Sady, 2 x Zady, 2 x Żydy

Certainly quite a lot to search for a lost package.

[0]: https://dane.gov.pl/pl/dataset/188,wykaz-urzedowych-nazw-mie...

Re: Charset="WTF-8"

#63

I have an 'æ' in my middle name (formally secondary first name because history reasons). Usually I just don't use it, but it's always funny when a payment form instructs me to write my full name exactly as written on my credit card, and then goes on to tell me my name is invalid.

Still much better when it fails at the first step. I once got myself in a bit of a struggle with Windows 10 by using "ł" as part of Windows username. Amusingly/irritatingly large number of applications, even some of Microsoft's own ones, could not cope with that.

Re: Charset="WTF-8"

#64
post #26

Earlier quoted context omitted.

> I'm curious if anyone can think of any other non-alphabetic characters used in legal names around the world, in other scripts? Latin characters are NOT allowed in official names for Japanese citizens. It must be written in Japanese characters only. For foreigners living in Japan it's quite frequent to end up in a situation where their official name in Latin does not pass the validation rules of many forms online. I…

Very interesting about Japan! To be clear, I wasn't thinking about within a specific country though. More like, what is the set of all characters that are allowed in legal names across the world? You know, to eliminate things like emoji, mathematical symbols, and so forth.

Ah, I see.

I don't know, but I would bet that the sum of all corner cases and exceptions in the world would make it pretty hard to confidently eliminate any "obvious" characters.

From a technical standpoint, unicode emojis are probably safe to exclude, but on the other hand, some scripts like Chinese characters are fundamentally pictograms, which is semantically not so different than an emoji.

Maybe after centuries of evolution we will end up with a legit universal language based on emojis, and people named with it.

Re: Charset="WTF-8"

#65
post #18

How do I allow "stępień" while detecting Zalgo-isms?

There's nothing special about "Stępień", it has no combining characters, just the usual diacritics that have their own codepoints in Basic Multilingual Plane (U+0119 and U+0144). I bet there are some names out there that would make it harder, but this isn't one.

Re: Charset="WTF-8"

#66
post #46

I have an 'æ' in my middle name (formally secondary first name because history reasons). Usually I just don't use it, but it's always funny when a payment form instructs me to write my full name exactly as written on my credit card, and then goes on to tell me my name is invalid.

As you may be aware, the name field for credit card transactions is rarely verified (perhaps limited to North America, not sure). Often I’ll create a virtual credit card number and use a fake name, and virtually never have had a transaction declined. Even if they are more aggressively asking for a street address, giving just the house number often works. This isn’t a deep cover but gives a little bit of a anonymity f…

It's for when things go wrong. Same as with wire transfers. Nobody checks it unless there's a dispute.

Re: Charset="WTF-8"

#67
post #64

Earlier quoted context omitted.

Very interesting about Japan! To be clear, I wasn't thinking about within a specific country though. More like, what is the set of all characters that are allowed in legal names across the world? You know, to eliminate things like emoji, mathematical symbols, and so forth.

Ah, I see. I don't know, but I would bet that the sum of all corner cases and exceptions in the world would make it pretty hard to confidently eliminate any "obvious" characters. From a technical standpoint, unicode emojis are probably safe to exclude, but on the other hand, some scripts like Chinese characters are fundamentally pictograms, which is semantically not so different than an emoji. Maybe after centuries o…

Chinese characters are nothing like emoji. They are more akin to syllables. There is no semantic similarity to emoji at all, even if they were originally derived from pictorial representations.

And they belong to the {Alphabetic} Unicode class.

I'm mostly curious if Unicode character classes have already done all the hard work.

Re: Charset="WTF-8"

#68

Earlier quoted context omitted.

If you just use the {Alphabetic} Unicode character class (100K code points), together with a space, hyphen, and maybe comma, that might get you close. It includes diacritics. I'm curious if anyone can think of any other non-alphabetic characters used in legal names around the world, in other scripts? I wondered about numbers, but the most famous example of that has been overturned: "Originally named X Æ A-12, the chi…

דויד Smith (concatenated) will have an LTR control character in the middle

Oh that's interesting.

Is that a thing? I've never known of anyone whose legal name used two alphabets that didn't have any overlap in letters at all -- two completely different scripts.

Would a birth certificate allow that? Wouldn't you be expected to transliterate one of them?

Re: Charset="WTF-8"

#69
post #46

Earlier quoted context omitted.

As you may be aware, the name field for credit card transactions is rarely verified (perhaps limited to North America, not sure). Often I’ll create a virtual credit card number and use a fake name, and virtually never have had a transaction declined. Even if they are more aggressively asking for a street address, giving just the house number often works. This isn’t a deep cover but gives a little bit of a anonymity f…

It's for when things go wrong. Same as with wire transfers. Nobody checks it unless there's a dispute.

The thing is though that payment networks do in fact do instant verification and it is interesting what gets verified and when. At gas stations it is very common to ask for a zip code (again US), and this is verified immediately to allow the transaction to proceed. I’ve found that when a street address is asked for there is some verification and often a match on the house number is sufficient. Zip codes are verified almost always, names pretty much never. This likely has something to do with complexities behind “authorized users”.

Re: Charset="WTF-8"

#70
post #18

How do I allow "stępień" while detecting Zalgo-isms?

Zalgo is largely the result of abusing combining modifiers. Declare that any string with more than n combining modifiers in a row is invalid. n=1 is probably a reasonable falsehood to believe about names until someone points out that language X regularly has multiple combining modifiers in a row, at which point you can bump up N to somewhere around the maximum number of combining modifiers language X is likely to hav…

I can point out that Greek needs n=2: for accent and breathing.
Post reply on HN