Live data from Hacker News

EuroLLM: LLM made in Europe built to support all 24 official EU languages

eurollm.io

61–70 of 629 posts

Re: EuroLLM: LLM made in Europe built to support all 24 official EU languages

#61
post #44

As expected, Europe finally catches up to 2024 and launches an LLM that barely competes against the heavyweights. The US and China are running rings around Europe. Mistral is an exception as it was funded by US VCs and they are a great example showing that without VC funding, Mistral would have been begging to the EU for a microsopic grant to train a LLM worse than Llama.

less exposure to a technology that doesn't bring that much revenue and it's not projected to do so in the upcoming years.

yep, Europe is demonstrating the same sort of strategic thinking that economic behemoths like the Smithsonian use

Re: EuroLLM: LLM made in Europe built to support all 24 official EU languages

#63
post #33

Earlier quoted context omitted.

Yup, most of Eastern Europe are Balto-Slavic. While the division from the Eastern Slavic languages (Russian, Belarussian, Ukranian, etc) is distant, they are still Slavic. From Eastern Europe, only Estonian is not a Slavic language.

Hungarian too, although there’s a question about whether Hungary is Eastern or Central Europe.

“There’s a question” implies that there is a ground truth that might be discovered to resolve this rather than simply a clash of different purely arbitrary definitions of the same terms.

Re: EuroLLM: LLM made in Europe built to support all 24 official EU languages

#64

How does this work? It seems like it, in most ways, it would be bad to train on 24 separate languages. That's just 24 partitions to the data. Seems really inefficient and better to simply train in the biggest (english) and translate. I do think this will introduce some biases that correlate with the English language. It would be interesting to see more specifically what this means. But regardless, I don't think you c…

If you train a model on multiple languages, you can use the model itself for translation. As well as allowing the model to naturally respond in the user's language.

Re: EuroLLM: LLM made in Europe built to support all 24 official EU languages

#65

1. It's a nice start, but the EU has to scale to Manhattan Project levels in order to properly compete with the US and China. 2. A credible scale effort for EU own silicon for AI Compute, wouldn't hurt either. 3. And this can only be achieved by vertical integration to combat fragmentation.

Good to distinguish between publicly funded research models (like this one) and commercial ones (like Mistral in France). What are the chinese and usa public research models like?

Re: EuroLLM: LLM made in Europe built to support all 24 official EU languages

#67
post #43
post #35

Earlier quoted context omitted.

Tomorrow there are elections in the Netherlands, and two parties are proposing adding Frysian to that list: https://neerlandistiek.nl/2025/10/kies-voor-taal/ Best get to retraining those models.

Each EU country nominates one official language for the EU, otherwise we'd have Catalan, Breton, Kashubian and many more.

They could get Austria to do it, as it presumably has a spare slot.

Re: EuroLLM: LLM made in Europe built to support all 24 official EU languages

#69
post #44

As expected, Europe finally catches up to 2024 and launches an LLM that barely competes against the heavyweights. The US and China are running rings around Europe. Mistral is an exception as it was funded by US VCs and they are a great example showing that without VC funding, Mistral would have been begging to the EU for a microsopic grant to train a LLM worse than Llama.

less exposure to a technology that doesn't bring that much revenue and it's not projected to do so in the upcoming years.

Why wasting money on trying to compete at all then?

Re: EuroLLM: LLM made in Europe built to support all 24 official EU languages

#70

How does this work? It seems like it, in most ways, it would be bad to train on 24 separate languages. That's just 24 partitions to the data. Seems really inefficient and better to simply train in the biggest (english) and translate. I do think this will introduce some biases that correlate with the English language. It would be interesting to see more specifically what this means. But regardless, I don't think you c…

nah, it's better to train on all languages. 24 partitions? you are gravely underestimating these models and how they represent things in their latents... transfers easily
Post reply on HN