Live data from Hacker News

Unicode 13.0

unicode.org

31–40 of 96 posts

Re: Unicode 13.0

#31

> Unicode 13.0 adds 5,930 characters, for a total of 143,859 characters. These additions include 4 new scripts, for a total of 154 scripts, as well as 55 new emoji characters. So how far off is Unicode from being 'done'? At what point will they be able to stop adding characters and scripts?

Until there's a kumquat emoji then it will not be done:

* https://www.emojis.com/food/fruit/

There will always be another Emoji that someone, somewhere wants to add.

Re: Unicode 13.0

#32
post #21

> Unicode 13.0 adds 5,930 characters, for a total of 143,859 characters. These additions include 4 new scripts, for a total of 154 scripts, as well as 55 new emoji characters. So how far off is Unicode from being 'done'? At what point will they be able to stop adding characters and scripts?

German, a European language that has been more or less standardized for several centuries, with a Latin-based alphabet, added a new letter (ẞ) to its alphabet in 2017. As long as that continues to happen, Unicode will have to add new characters, even if no more ancient scripts are discovered and no new writing systems are developed for currently unwritten languages.

https://en.wikipedia.org/wiki/Capital_%E1%BA%9E

That letter, the capital eszett, has existed in German typefaces since at least 1905:

> Historical typefaces offering a capitalized eszett mostly date to the time between 1905 and 1930. The first known typefaces to include capital eszett were produced by the Schelter & Giesecke foundry in Leipzig, in 1905/06. Schelter & Giesecke at the time widely advocated the use of this type, but its use remained very limited.

Eszett is usually just a lower-case form; it is most often upper-cased to SS, or two capital letter esses, being an example of how changing case does not always preserve the number of letters in a text. Capital eszett is very rare, but it was in uncommon usage in German text, and so it was added to Unicode.

Re: Unicode 13.0

#33
post #26
post #4

Earlier quoted context omitted.

Once people stop needing and/or inventing new characters and scripts.

either that or when there are 2^24 used codepoints.

The available space is closer to 2^20 (0-10FFFF, minus surrogate pairs, depending on whether you are talking about Unicode scalar values or code points).

Re: Unicode 13.0

#35
post #11

Earlier quoted context omitted.

I suspect emojis as characters will soon be replaced with graphics. Consumers are annoyed or confused when emojis look slightly different on different platforms.

I will regret this but... Animated ligature emojis or some sorta crontraption... Animated characters is the next stage. If emojis are different between carriers meh. Course I find it silly that they swapped a real gun emoji for a water gun. How will my friends know I intend to go hunting and not play water guns with my nephews or whatever.

Uhh...perhaps try using actual words if you want to make your meaning clear? After all, if you used a "real gun" emoji, your friends might just think you were planning a massacre...

Re: Unicode 13.0

#36

Earlier quoted context omitted.

I suspect the latter is a significantly smaller group though - I would expect people care more about their messages coming across unchanged

Interesting. I’d think the opposite: I bet the number of times the average user looks at their own message on someone else’s device is pretty low, while they might have a chance to notice the differences in emoji between app A and app B on their own device several times per day.

> I bet the number of times the average user looks at their own message on someone else’s device is pretty low

That is not the issue. The issue is that users expect the emoji they send to look the same to the receiver.

Geeks know that emojis are "characters" and therefor might look slightly different in different fonts. But this is not how the average user sees it. They see emojis as small images and expect them to transfer unaltered just like any other image they put into a message.

Re: Unicode 13.0

#37
post #21

Earlier quoted context omitted.

German, a European language that has been more or less standardized for several centuries, with a Latin-based alphabet, added a new letter (ẞ) to its alphabet in 2017. As long as that continues to happen, Unicode will have to add new characters, even if no more ancient scripts are discovered and no new writing systems are developed for currently unwritten languages.

Small precision for those who don't know the context: the Eszett (which comes from the ligature of 'ss') existed for centuries already in German writing. 2017 is just the date of its official integration in the alphabet, so it's not a 'new' letter created from scratch. I say that because I remember learning it at school a few decades ago (even if at the time we were warned the subject was touchy), and I was surprised…

The Eszett (ß) has been already standardized for several decades (since 1986: with ISO 8859-1 aka latin-1).

The newly added letter is the "capital letter Eszett", which did not exist until recently. One could argue that this new letter is not really needed, as Eszett does not appear in capitalized form except when a word is in all-caps, and was then simply written as "SS".

Re: Unicode 13.0

#38
post #21

Earlier quoted context omitted.

German, a European language that has been more or less standardized for several centuries, with a Latin-based alphabet, added a new letter (ẞ) to its alphabet in 2017. As long as that continues to happen, Unicode will have to add new characters, even if no more ancient scripts are discovered and no new writing systems are developed for currently unwritten languages.

Small precision for those who don't know the context: the Eszett (which comes from the ligature of 'ss') existed for centuries already in German writing. 2017 is just the date of its official integration in the alphabet, so it's not a 'new' letter created from scratch. I say that because I remember learning it at school a few decades ago (even if at the time we were warned the subject was touchy), and I was surprised…

And some more precision: this is about the uppercase Eszett. The lowercase variant has existed for ages, and before 2017 was officially uppercased into SS or SZ.

Re: Unicode 13.0

#39
post #11

Earlier quoted context omitted.

I suspect emojis as characters will soon be replaced with graphics. Consumers are annoyed or confused when emojis look slightly different on different platforms.

I'm confused. How would this solve the problem? Whilst looking different is clearly an issue (e.g. the gun being depicted a water gun), I'm not sure what your solution would look like? Images hosted somewhere? You'd get consistency, but you'd lose universality. It goes without saying, but you can use Emoji practically _anywhere_ that support text. That's huge.

Although it depends on having the same level of emoji support across all the relevant systems/devices. If you send a message with Unicode 13 emoji to me and I read it on my phone that's stuck on Unicode 10, it's not much use. Whereas if these silly little images were represented by links to a canonical repository somewhere, newly-added images could automatically work even in pre-existing systems.

On the other hand, that'd be dependent on connectivity.

Re: Unicode 13.0

#40

Earlier quoted context omitted.

And emojis, don't forget about the all-important emojis.

You are probably joking, but I am seriously annoyed there's no donkey emoji.

My wife and I call each other "donkey", and some years ago we used the horse emoji, which was low resolution enough to look as a donkey if you squinted. But modern emojis are too high resolution and it definitely looks like a horse now. So I feel your pain.
Post reply on HN