Live data from Hacker News

Don't guess my language

vitonsky.net

361–370 of 392 posts

Re: Don't guess my language

#361
post #351

Earlier quoted context omitted.

We can find a sorting order with the minimal total distance between where we place a language entry and where this entry would be in that language. If there’s no pair of languages A and Ä such that A > Ä in one and Ä > A in the other, then (I guess???) this total distance will be zero.

> A and Ä Coincidentally, the expected position of "Ä" can vary wildly. Is it an umlauted A, normalized as AE, or a distinct letter coming after Z?

> "Ä"

OT, but this looks like an adorably blushing hen to me

Re: Don't guess my language

#362
post #160
post #87

Earlier quoted context omitted.

duckduckgo seems to do it as well

Yep, this is coming from Reddit itself. It's using different URLs, and they seem to be making an effort to SEO-rank those translations.

It's interesting that the quality is so low. You can do very good translations for many languages today even with fairly cheap LMs, but for some reason (cost?) automated translation online seems to be still mostly at Google Translate level.

Re: Don't guess my language

#363
post #48

Earlier quoted context omitted.

> I suspect the idea that a person speaks more than one language is absent in US silicon valley. Which has been baffling to me considering how many foreigners work at these companies.

I think it's more a matter of "why would they have their system language set to X if they speak Y? If they want Y, they should just set their system language to Y!" It's the idea that the user has a preference for something, and it applies always and everywhere, even when it's not applicable.

This is actually something that foreigners working at Big Tech US companies should be able to understand very well, because English as a system language is often how software developers set things up for themselves regardless of their native language.

But they don't make those decisions. It's a UX thing, which means that in practice whoever is in charge of "driving up the numbers" is going to be making the decision; the engineers just get to cuss while implementing it.

Re: Don't guess my language

#364

Much more importantly: Never ever auto-translate content to the user's language. Present what languages you actually have the data in. The user is smart enough to click the "translate" button in the web browser should they want. That translation is also likely to be better quality. English is not my first language. Or my second. But I understand it well enough to work in it every day. And I never ever want to wade th…

Machine translation has been good enough for a few years, that even native speakers don't notice it.

If done with SOTA LLMs, yes, depending on the language (although even then I would dispute the "even native speakers don't notice it" part).

But SOTA LLMs cost money, so what you usually get in practice for auto-translation at scale (like YouTube) is around Google Translate level, which is hardly good enough.

Re: Don't guess my language

#365

Earlier quoted context omitted.

Machine translation has been good enough for a few years, that even native speakers don't notice it.

If done with SOTA LLMs, yes, depending on the language (although even then I would dispute the "even native speakers don't notice it" part). But SOTA LLMs cost money, so what you usually get in practice for auto-translation at scale (like YouTube) is around Google Translate level, which is hardly good enough.

I don't know what level you consider DeepL or Kagi Translate, but those two can translate at a level which is indistinguishable from a native speaker.

Re: Don't guess my language

#366

Earlier quoted context omitted.

If done with SOTA LLMs, yes, depending on the language (although even then I would dispute the "even native speakers don't notice it" part). But SOTA LLMs cost money, so what you usually get in practice for auto-translation at scale (like YouTube) is around Google Translate level, which is hardly good enough.

I don't know what level you consider DeepL or Kagi Translate, but those two can translate at a level which is indistinguishable from a native speaker.

These types of statements are hard to take seriously. Those tools are remarkably good for what they are, and the fact that they can help me understand texts in foreign languages is amazing. But indistinguishable from a professional native translator (which I take it you mean, being a language native is not enough to be a translator) is something else entirely.

If you really believe that, why not recruit a professional, do a blinded study, and if it's really industinguishable you can probably get a nice publication somewhere. AI companies are flush with cash and will not think twice about sponsoring such a study.

Re: Don't guess my language

#367

Earlier quoted context omitted.

The Wikipedia sort for the languages is as I stated above, with Literary Chinese and Japanese between Wu Chinese and Yue Chinese. I explained why it was sorted that way, because radical is considered first. You could not explain why Japanese appeared between Wu and Yue because you insisted and continue to insist that radicals are not used. I didn't say sorting is never done by stroke count alone. But I have seen radi…

> You could not explain why Japanese appeared between Wu and Yue because you insisted and continue to insist that radicals are not used. I can't explain that because it's part of a different logical group, with its name written in a different script.† This puts it parallel to the Chinese options and to Korean. > The Wikipedia sort for the languages is as I stated above I took you to be describing the sort order for c…

  > all of the languages mentioned so far appear before the "Latin alphabet" 
  > style languages, but 閩南語 and 閩東語 appear after them.
Could it have something to do with Minnan and Mindong Chinese articles being written in a Latin script, (despite the language name showing in both Chinese characters and Latin letters) ?

Re: Don't guess my language

#368

Earlier quoted context omitted.

If done with SOTA LLMs, yes, depending on the language (although even then I would dispute the "even native speakers don't notice it" part). But SOTA LLMs cost money, so what you usually get in practice for auto-translation at scale (like YouTube) is around Google Translate level, which is hardly good enough.

I don't know what level you consider DeepL or Kagi Translate, but those two can translate at a level which is indistinguishable from a native speaker.

This is emphatically not true in case of DeepL in my experience. I haven't tried Kagi but I wouldn't be surprised if they are using the same LLM that powers their assistant for it.

Also, this kind of claim is meaningless unless you specify the language being translated, because quality varies widely depending on that, even across popular languages, never mind small regional ones. If you try to translate, say, to Chechen using GPT-o3 or Gemini 2.5 Pro, don't expect that translation to be accurate or even grammatically correct.

Re: Don't guess my language

#369
post #9

My biggest annoyance is with Google. They know who I am, they know I am traveling, they know my language preferences (English) and yet I still get language based on my location on certain pages. I let you track me Google, please use it for some good UX and not just advertising.

About a decade ago Google decided that all maps in Finland should have the street names in Swedish.

Which is kinda valid, in the southern and south-western parts this is done because there is a significant Swedish-speaking minority so most cities and streets have names in both languages.

But at the time I lived in central Finland, where the streets DIDN'T have official Swedish names, they just ... translated them. Which was super fun for navigating.

Re: Don't guess my language

#370

Earlier quoted context omitted.

> You could not explain why Japanese appeared between Wu and Yue because you insisted and continue to insist that radicals are not used. I can't explain that because it's part of a different logical group, with its name written in a different script.† This puts it parallel to the Chinese options and to Korean. > The Wikipedia sort for the languages is as I stated above I took you to be describing the sort order for c…

> all of the languages mentioned so far appear before the "Latin alphabet" > style languages, but 閩南語 and 閩東語 appear after them. Could it have something to do with Minnan and Mindong Chinese articles being written in a Latin script, (despite the language name showing in both Chinese characters and Latin letters) ?

As far as I know, sure, it could.
Post reply on HN