Live data from Hacker News

Announcing Unicode 7.0

unicode-inc.blogspot.com

71–80 of 88 posts

Re: Announcing Unicode 7.0

#71
post #65
post #61

Earlier quoted context omitted.

But 粪 is also a character and means poo as well. It might not look like poo to you, but there's also examples of pictograms in the same script: 田 means field and looks like one. Of course these are "official part of a language". But why? Under whose authority? You could cite "actual usage", but many of the emoji in Unicode certainly have seen actual usage for 10+ years in East Asia, so they qualify. And many of the l…

Oh, the irony: http://oi60.tinypic.com/21o111t.jpg That's pretty much my point.

So you're using Windows XP or something similarly mired in the 90s which doesn't have fonts installed with glyphs for a standard character. That's bad for you and a small number of other people but that hardly says it's not useful to the billion or so people who can read Chinese.

The nice part about having it Unicode is that if you copy and paste that character somewhere else, it'll still work. It never goes through what we used to have to deal where the same character value couldn't reliably be processed without knowing the intended encoding with certainty.

Re: Announcing Unicode 7.0

#72
post #69

Earlier quoted context omitted.

> I'm getting far from the original topic with this rambling, but it still reads on it in one sense: There's a lot of variety to writing systems, and it's worth thinking about emoji as a spot on the spectrum instead of in isolation. Except that emoji aren't a writing system. Writing systems encode language. Emoji encode emotions, perhaps, or suggest collections of ideas/emotions, but don't convey language per se beca…

Are you suggesting we turn falling into a specific set of relationships between symbols an entry requirement for the Unicode database, though? How would that look like in practice? I agree you're highlighting a useful difference, but I'm not sure about implications.

Nah, I think keeping unicode as "characters" is fine. It's perfectly reasonable to encode non-linguistic data as characters if they're using in such wise.

Edit: to circle back, you said:

  How are emoji different from other logograms?
and everything thereafter was pointless pedantry on my part. Sorry for wasting valuable minutes of your life.

Re: Announcing Unicode 7.0

#73
post #69

Earlier quoted context omitted.

Are you suggesting we turn falling into a specific set of relationships between symbols an entry requirement for the Unicode database, though? How would that look like in practice? I agree you're highlighting a useful difference, but I'm not sure about implications.

Nah, I think keeping unicode as "characters" is fine. It's perfectly reasonable to encode non-linguistic data as characters if they're using in such wise. Edit: to circle back, you said: How are emoji different from other logograms? and everything thereafter was pointless pedantry on my part. Sorry for wasting valuable minutes of your life.

No waste at all, you raised a good point :). Emoji are technically ideograms, not logograms ... though some of the Han characters are ideographic in origin/conception too (and then you get into cool things like compound ideograms), they just didn't stay purely ideographic.

Re: Announcing Unicode 7.0

#74
post #71
post #65

Earlier quoted context omitted.

Oh, the irony: http://oi60.tinypic.com/21o111t.jpg That's pretty much my point.

So you're using Windows XP or something similarly mired in the 90s which doesn't have fonts installed with glyphs for a standard character. That's bad for you and a small number of other people but that hardly says it's not useful to the billion or so people who can read Chinese. The nice part about having it Unicode is that if you copy and paste that character somewhere else, it'll still work. It never goes through…

No, Chrome, Debian testing.

Re: Announcing Unicode 7.0

#75

Earlier quoted context omitted.

I had the exact same question pop up in my mind after seeing the emoji's. I understand all characters can be considered nothing but pictures, but a thermometer, really? Is this used in any widespread language to be of significance to Unicode? Is Unicode trying to also offer some kind of clipart capability? I think we've seen our fair share of derailed projects in IT. Unicode is certainly a very critical project, whic…

Is the thermometer symbol used widely? No idea. Is emoji used widely? Yes, In Japan 🇯🇵 (and Korea? 🇰🇷) and has been for > 15 years. A common email/msg composed by a Japanese woman 👧 age 15 to 35 uses many emoji. I realize this message will not render correctly in Chrome because Chrome, at least Desktop Chrome, does not yet support emoji 😭

Is the problem with Chrome or the OS and font selection? I think everything renders on my Win 7 + FX 30 setup I'm using now. Generally, I have the best success with Win 8 + IE 11.

Re: Announcing Unicode 7.0

#76
post #67
post #65

Earlier quoted context omitted.

Oh, the irony: http://oi60.tinypic.com/21o111t.jpg That's pretty much my point.

For the record, the example was a Chinese character, which qualifies as "normal" according to the rulebook you set up. It's in active use in Mandarin and Japanese, and I'd think recognizable to many others (e.g. Cantonese readers or Koreans and Vietnamese). And it'll certainly show just fine on the default installs of most systems on the market right now. http://www.unicode.org/cgi-bin/GetUnihanData.pl?codepoint=7C..…

The point is that it will never happen that all codepoints are available everywhere (every font, evers OS, every device), and i certainly don't want a gazillion of pictures loaded on my smartphones little memory, just because some people think it necessary to have a symbol for poop in fonts which is never used anyway. I even consider it a total waste in my laptops 16GB of memory, to be honest. It's not a waste when it serves the real purpose of a language. A chinese word for poop? Sure. A picture of actual poop? No. And where does it stop? Can i have some dog poop and cat poop codepoints please? I'm sure elephant poop looks very different to cat poop, i want a code point for that too.

I say it again: If you want pictures and cliparts, use .png. If you want to write actual text, use fonts. It's as simple as that.

Re: Announcing Unicode 7.0

#77
post #54
post #43

Earlier quoted context omitted.

And you consider Wingdings a good thing in the first place?! It's an abomination that should never have happened.

Never said that, but the current situation is that we have glyphs that people are using in documents that can't be represented in Unicode, and that's pretty much the main reason there are Unicode revisions in the first place.

Why would documents using Webdings stop working today? I'm pretty sure they still work. That's a lame excuse. Considering your point, wouldn't it have been the chance to clean up this mess with unicode instead of piling up shitty icons in a standard?

Re: Announcing Unicode 7.0

#78
post #70
post #12

Do we really need small pictures in fonts?

You might not need them but consider this for a moment: When I was a kid, there was a boy in my class who was mute. He carried a chalkboard and flashcards to facilitate communication. Eventually, he was chosen to pilot a device that was sort of like a keyboard with oversized keys. Each key had a pictograph on it[0]. Wouldn't it be nice for him if the flashcards and keyboard device had a consistent set of characters t…

And that device would have a keyboard with thousands of keys?

Btw. that device exists today, and it's called a tablet. I'm pretty sure i have seen stuff exactly like that and it serves a great purpose.

The consistent set of "characters" (pictures in this case) would still look very different on each device or application that uses another font. Because a banana in one font could look very different in another font. The only thing that actually would be consistent if the content he uses would use actual pictures! (A website displaying a banana on his touchpad can look different on his laptop, can't it?).

Re: Announcing Unicode 7.0

#79
Everybody's talking about new glyphs/codepoints, but unicode is actually way more than a list of codepoints.

Note the announcement also includes:

* Changes to the unicode collation algorithm for locale-dependent sorting (I haven't figured out what htey are yet)

* New character properties and values, used for case-changing and line-breaking behavior (i think line-breaking might be locale-dependent too).

* some other stuff

The unicode algorithms for _doing things_ with text in appropriate ways are actually way more mind-blowing to me than just the directory of glyphs. When you start thinking about it and get into the weeds and realize how other languages work very differently than English with regard to some of this stuff -- proper "uppercase" or "sort" or "are these two strings semantically the same" across the global universe of glpyhs -- is _really_ challenging, and unicode does a pretty amazing job of it.

It's sadly way more confusing to figure out if/when the unicode library of your choice has been updated to use unicode 7.0 (or even 6.0) text algorithms and full character database with properties, rather than to figure out if a given font includes the new glpyh or whatever.

Post reply on HN