Live data from Hacker News

How a comment on Hacker News led to 4½ new Unicode characters

unicodepowersymbol.com

421–429 of 429 posts

Re: How a comment on Hacker News led to 4½ new Unicode characters

#421
post #378

Earlier quoted context omitted.

There's another option: 3) People who love colourful images will use stickers in Facebook Messenger, LINE, Viber, and soon iMessage. I'm sure WeChat has them too. It's basically like 2), except we've moved from proprietary codepoints to proprietary protocols. I don't mind characters like and or even good old ︎ (which has always been too tiny for its own good). These work in black and white, in different artistic styl…

Interesting. Did you try to include some emoji in your comment? They did not get included: > characters like and or even good old ︎ (which

Oops, thanks. Well, that explains why I've never seen Emoji on Hacker News. And I've missed the edit window, so I can't fix my post.

Should have been:

> I don't mind characters like ((yellow smiling Emoji)) and ((thumbs up Emoji)) or even good old ︎((pre-Emoji Unicode smiley)) (which has always been too tiny for its own good). These work in black and white, in different artistic styles, and they're a fairly limited set.

> But now we're going down the road where we get new stuff like tacos and unicorns every year. And even though Unicode is an industry standard, the pictures need to look like Apple's bitmaps to avoid confusion, and the Unicode standard changes so often that you have to manually keep track of who can already see ((upside-down smiling Emoji)) and whose computer/phone/browser/messenger software is too old.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#422
post #372

Earlier quoted context omitted.

It is true for lots of characters (so I guess I was being a little hyperbolic when I said "all"), and you cannot rely on choosing the correct code points in order to have a text display Japanese or Chinese. You need to tell your rendering program (often through choice of font) if things are to be rendered with Japanese or Chinese forms. I wouldn't know how to show you examples here, as 直 will 直 display the same since…

Aren't they putting the disunified characters into the U+2xxxx plane now? Han unification is generally seen as a bad choice in retrospect, but it was something Unicode had to do when it looked like 2^16 codepoints were all they were going to get.

Never heard of that, but I would appreciate if all the characters with different glyphs had different codepoints. Do you have a source? Do you know what happens to the "unified" code-points?

Re: How a comment on Hacker News led to 4½ new Unicode characters

#423
post #302
post #279

Earlier quoted context omitted.

> there is no common fallback font that is used by all systems (win, osx, linux, android ...) as a default. Again, meaningless. Why should there have to be a common fallback font?

to make sure that even the most idiosyncratic choice of characters can be displayed everywhere according to the intention of the person who picked them.

Unicode supplies sample glyphs for all characters. That's plenty.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#424
post #405

Earlier quoted context omitted.

The character 囍 appears in other standards (big5, JIS, ISO2022_JP), so Unicode is basically obligated to include it for compatibility. The running text requirement is for symbols, so it doesn't apply to Chinese: http://www.unicode.org/pending/symbol-guidelines.html

As to compatibility, excellent point. The word "running" doesn't appear on that page. (Actually, no requirements at all appear on that page; it speaks strictly in terms of strengthening or weakening the case for inclusion, not disqualifying.) Can you explain briefly why that page is evidence that the running text requirement does not apply to Chinese, and where it specifies what the running text requirement is? Alter…

I thought the symbol guideline page discussed "running text", but I guess not. Apparently the "running text" requirement isn't part of the published criteria even though it is enforced in discussion.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#425
post #363

Earlier quoted context omitted.

> Are you saying that a company with voting rights shouldn't be allowed to have influence on what goes into Unicode? One might have a right to do something and yet be wrong to do it. Apple had every right to do what they did, but they were completely, totally and undeniably in the wrong to do it. Everyone associated with their action should be ashamed. Honestly, they should all resign: their behaviour demonstrates th…

Your comment is ridiculously extreme. They were not "undeniably" in the wrong. You think they were wrong, but that is an highly subjective opinion. In fact, I don't even agree that they were wrong to do this at all. I think it's perfectly reasonable to argue against the inclusion of more gun imagery in Unicode. Also, if you think Apple was wrong, you must also think that Microsoft was (they voiced support), and every…

Just like how North Korea isn't "undeniably" wrong in censoring speech to such an extreme. This is a very 1984-esque "solution" to a problem -- don't want to acknowledge positive use of guns? Good news, we can just erase them from our language! The way we treat emoji has some very serious similarities to Newspeak.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#426

Earlier quoted context omitted.

Honestly, I think "SVG over UTF" makes a lot more sense. It's impossible to make a character set that supports every character known to man, because that just adds undue effort on every computer maker, ect, to keep up. So why don't we pick a very good set: perhaps every letter in every language in common use for the past 200 years? Then, for the oddball symbols that someone wants to mix in text, there can be some kin…

Because it's easier to throw in random icons than to actually accomplish the goal of "every letter in every language in common use for the past 200 years", or even "past 20 years". Or, put another way: 'We have an unambiguous, cross-platform way to represent “PILE OF POO” (), while we’re still debating which of the 1.2 billion native Chinese speakers deserve to spell their own names correctly.' https://modelviewcultu…

Correct me if I'm wrong, but isn't the Han Unification project more about unifying semantically distinct, but visually identical characters under the same codepoint (rather than grouping together similar-looking codepoints as the article suggests)? As far as I'm aware it's more along the lines of reusing the codepoint for 'a' when encoding both English and Spanish text. Am I mistaken in thinking this?

Re: How a comment on Hacker News led to 4½ new Unicode characters

#427
post #425

Earlier quoted context omitted.

Your comment is ridiculously extreme. They were not "undeniably" in the wrong. You think they were wrong, but that is an highly subjective opinion. In fact, I don't even agree that they were wrong to do this at all. I think it's perfectly reasonable to argue against the inclusion of more gun imagery in Unicode. Also, if you think Apple was wrong, you must also think that Microsoft was (they voiced support), and every…

Just like how North Korea isn't "undeniably" wrong in censoring speech to such an extreme. This is a very 1984-esque "solution" to a problem -- don't want to acknowledge positive use of guns? Good news, we can just erase them from our language! The way we treat emoji has some very serious similarities to Newspeak.

That's absurd. Nobody's censoring anything here. Apple not wanting to add a new gun emoji is in no way preventing you from talking about guns. Emoji isn't a replacement for English and nobody is forcing you to "speak" in all emoji.

Re: How a comment on Hacker News led to 4½ new Unicode characters

#428

How long before we need a defusedunicode to protect users and programs from confusion and scams? https://pypi.python.org/pypi/defusedxml

Actually there is something like that in the internationalised DNS standards (that is, internet domain names that look like .xn--xyz-abc sometimes and .日本 other times.) There's a blacklist of certain Unicode characters that are disallowed in domain names because they resemble more commonly used characters. See https://en.wikipedia.org/wiki/IDN_homograph_attack

Re: How a comment on Hacker News led to 4½ new Unicode characters

#429
post #108
post #69

Earlier quoted context omitted.

Yes, but they have alternate italic forms, for example. Sure, some one of the glyphs like с doesn't have an alternate italic form. Since the other ones do, it would be weird to only assign a separate codepoint to some of them and overlap the others. It would be a workable solution, but still weird.

You mean like Han unification?

Which is a bad idea because the characters don't look right unless you use a Japanese font. But if you want to write an article comparing Japanese and Chinese characters, you have to use two different fonts.
Post reply on HN