Live data from Hacker News

Help! Is This Arabic?

isthisarabic.com

71–80 of 304 posts

Re: Help! Is This Arabic?

#72

Crazy how there’s “two billion people” who can read Arabic and seemingly 0 of them are making serious contributions to Arabic text rendering except this one specific guy. Weird.

Two billion people can read it to some degree because of Islam. A minority of them actually use the script to write things daily.

And there are plenty of people other than this guy making contributions to it, you’ve only just read a singular post by a single person outlining his most common gripes with companies that do not consult native speakers for their translations. It’s not an exhaustive list of people working in the sphere. Imagine if you read a blog post by Linus Torvalds, and thought: ‘wow, it’s crazy that no one else is making contributions to this cool open source kernel?’ - that’s exactly what you’ve done in your comment.

Re: Help! Is This Arabic?

#73
post #56

This website is well-meaning, but will be difficult to parse if you don't have elementary Arabic and can't tell apart ال from ل ا. What it really needs is a simple reference input string, examples of how it gets broken, and what to do to fix them. The middle case, where the sentence is correctly rendered RTL but the individual words are LTR (breaking the ligatures), is particularly common and insidious because it loo…

> a simple reference input string, examples of how it gets broken, and what to do to fix them. I really like this idea. Just have a standard set of strings covering all edge cases (even the sprawling labyrinth that is bidi) with a visual reference that shows how the correct rendering of each string would look like. Each entry would also have a description of the problem and suggested solutions. Unlike the solutions i…

I'm not sure I understand the idea. It sounds insane to me because I feel like there's probably trillions of combinations (and it would be insane to expect to be able to cover every specific example of incorrect text), and I thought the website was pretty clear and provided good examples.

Re: Help! Is This Arabic?

#74
post #31

Also add "use the correct Arabic for your target audience". "Arabic" is not monolithic and using the wrong dialect can go over as well as using Scots on a web page intended for Jamaicans.

As an Arabic speaker, baby steps! For spoken Arabic, a decent placeholder is either MSA (standard Arabic) or a generally understood dialect (Egyptian or Levantine). The latter is what a lot of the recent Arabic game dubs have been doing (e.g., Ubisoft). Another example: Pixar and Disney use Egyptian Arabic for their dubs. Edit: Oh, and for written Arabic, you should almost always be using MSA. That is, you shouldn’t…

I pity the poor dev who has a Moroccan co-worker that helpfully offers to translate their application into darija thinking it’s Arabic

Re: Help! Is This Arabic?

#75
post #62
post #19

Earlier quoted context omitted.

It's weird, because the number is off by an order of magnitude. "If you count all of the varieties of today’s Arabic together, you can safely estimate that there are about 313 million Arabic speakers in the whole world, making it the fifth most-spoken language globally behind Mandarin, Spanish, English and Hindi." [1] [1] https://www.babbel.com/en/magazine/how-many-people-speak-ara...

The writer most likely count the moslem population. Most moslem don't speak arabic but they recite quran (written in arabic) daily. Most of them will notice if a ال written as ل ا

Egypt has ~100 million people, the vast majority of whom are going to be native Arabic speakers. Algeria, Sudan, and Iraq also have ~40 million people each, again, most of whom are native Arabic speakers (of different varieties). Morocco, Saudi Arabia, and Yemen are no slouches either. In short 300 million native speakers of various Arabic languages is extremely plausible.

FWIW, the usual estimate for worldwide adherents to Islam is ~1 billion.

Re: Help! Is This Arabic?

#76

Earlier quoted context omitted.

Yup, that's the exact point they're making. Non-Arabic-speaker sees Arabic script and assumes it's the Arabic language, so presents it to users that don't understand it, such as presenting Farsi text to Egyptian Arabic speakers. Or, to take the Latin script example, presenting French text to English speakers.

Well as the Persian and Arabic alphabet are nearly identical, unicode doesn't appear to store two distinct copies, instead just storing one plus the additional characters for Persian. Oddly though, it stores several copies of identical chinese characters. Traditional, simplified, and the Japanese loaned characters kanji.

> Oddly though, it stores several copies of identical chinese characters. Traditional, simplified, and the Japanese loaned characters kanji.

Actually, all of those characters are usually unified into one code point. This is controversial enough on its own that there's a Wikipedia page on it: https://en.wikipedia.org/wiki/Han_unification

Re: Help! Is This Arabic?

#77
post #69
post #41

Earlier quoted context omitted.

I could understand if this were a non-alphabet system like Chinese where the information coding is that dense per character. But for an alphabet system, why the mismatch between the bits of information per character and the bits needed to visually represent that character? The Latin alphabet mostly fits on a 7 seg display, and 26 letters requires a theoretical minimum of 5 bits to encode, so the graphical efficiency…

Because the information coding is dense. Arabic script has a lot of tiny shapes that indicate what letter it is and it also has diacritic marks, which on their own are similarly intricate and small. You can't represent one line, two lines, a dot, a circle, a hamza, etc. above, below, to the side of particular characters, without a lot of resolution.

Another challenge for representing Arabic a display made of line segments is all the curves. If it weren't for the dots, I imagine you could write consonant-only Arabic on something close to a seven-segment display, although it would look bizarre because all the letters would be isolated and there would be so many more straight line segments than in normal Arabic script. (An absolute majority of the isolated forms of Arabic letters are made up only of curves, whereas an absolute majority of Latin capital letters are made up only of straight line segments!)

I wonder how Arabic ended up with characters that differ only by dot positioning (like ب ت ث, and also ن whose combining form is especially similar to the combining forms of those). By contrast, the most similar Latin characters might be EF / CG / IJ / MN / UV which I think are clearly more different than the Arabic characters that differ only by dot quantity and positioning.

(The Il1| are also famously confusable in some Latin fonts, but one could say having to worry about these comes late in the history of Latin script writing. Although the scribal minim https://en.wikipedia.org/wiki/Minim_(palaeography) used at some points to write u, i, m, and n is every bit as confusing as any Arabic character.)

Re: Help! Is This Arabic?

#78

Earlier quoted context omitted.

Yup, that's the exact point they're making. Non-Arabic-speaker sees Arabic script and assumes it's the Arabic language, so presents it to users that don't understand it, such as presenting Farsi text to Egyptian Arabic speakers. Or, to take the Latin script example, presenting French text to English speakers.

Well as the Persian and Arabic alphabet are nearly identical, unicode doesn't appear to store two distinct copies, instead just storing one plus the additional characters for Persian. Oddly though, it stores several copies of identical chinese characters. Traditional, simplified, and the Japanese loaned characters kanji.

> Traditional, simplified, and the Japanese loaned characters kanji.

Traditional, simplified, Japanese, Korean. Some are identical, some different. I wonder why these rendering details were not left to fonts and you can type 雨 and ⾬ for example. Maybe stronger presence of those countries in whatever SDOs and tech integration has to do with it.

Re: Help! Is This Arabic?

#79
> Photoshop breaks Arabic. We even have a website that will break our Arabic so that Photoshop breaks it back to normal. Yes, Adobe is aware. No, they do not care.

Photoshop since time immemorial had a Middle East edition (ME) which supported RTL and used to curb piracy

I don’t know how it’s with their new pseudo-sass but I think they integrated it into their main product

Re: Help! Is This Arabic?

#80
post #14

Earlier quoted context omitted.

I would be careful linking "Muslims" and "Arabic readers/speakers". I know Muslims who would be challenged to speak more than 3 words of Arabic, as well as Arabic speakers who are Christians or atheists.

i think every muslim can speaks more than 3 words, every muslim known what al fatiha is. that's 25 words.

I'm completely ignorant of Muslim culture but there was a time where Catholics recited prayers and liturgy in latin and plenty of practitioners recited the sounds without really understanding them.

Source: my grand aunts saying "arapreme" where "ora pro nobis" would go (the former kinda sounds like an Italian word). Also probably related, "hocus pocus" is the "magic" expression "hoc est corpus" ("this is the body [of Christ]").

Post reply on HN