I'll say it again: this is the consequence of Unicode trying to be a mix of html and docx, instead of a charset. It went too far for an average Joe DevGuy to understand how to deal with it, so he just selects a subset he can handle and bans everything else. HN does that too - special symbols simply get removed. Unicode screwed itself up completely. We wanted a common charset for things like latin, extlatin, cjk, cyri…
But hey, multiocular o! https://en.wikipedia.org/wiki/Cyrillic_O_variants#Multiocula... (TL;DR a bored scribe's doodle has a code point)
> The character was proposed for inclusion into Unicode in 2007 and incorporated as character U+A66E in Unicode version 5.1 (2008). The representative glyph had seven eyes and sat on the baseline. However, in 2021, following a tweet highlighting the character, it came to linguist Michael Everson's attention that the character in the 1429 manuscript was actually made up of ten eyes. After a 2022 proposal to change the character to reflect this, it was updated later that year for Unicode 15.0 to have ten eyes and to extend below the baseline. However, not all fonts support the ten-eyed variant as of November 2024.
So not only did they arbitrarily add a non-character (while arbitrarily combinding other real characters) but they didn't even get the glyph right and then changed it to actually match the arbitrary doodle.