Live data from Hacker News

Open Euro LLM: Open LLMs for Transparent AI in Europe

openeurollm.eu

1–10 of 289 posts

Re: Open Euro LLM: Open LLMs for Transparent AI in Europe

#3
> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence

As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be.

I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep relying on Google to do proper multilinguality in open models because Mistral can't be arsed to bother beyond French and German.

Re: Open Euro LLM: Open LLMs for Transparent AI in Europe

#4

> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep r…

>Mistral can't be arsed to bother beyond French and German.

Any more details here or a writeup you can link to?

Re: Open Euro LLM: Open LLMs for Transparent AI in Europe

#5
post #4

> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep r…

>Mistral can't be arsed to bother beyond French and German. Any more details here or a writeup you can link to?

My own experience mainly, only Gemma seems to have been any good for Slavic languages so far, and only the 27B when unquantized is reliable enough to be in any way usable.

Ravenwolf posts tests on his German benchmarks every so often in locallama and most models seem to do well enough, but I've heard some claims from people about Mistral's being their favorite models in German anyhow. And I think Mistral-Large scores higher than Llama-405B in French on lmsys and that's at least something one would expect from a French company.

Re: Open Euro LLM: Open LLMs for Transparent AI in Europe

#6

> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep r…

What do you mean by relying on Google?

Llama 3.1 and DeepSeek v3/R1 largest models are rather good at even a niche language like Finnish. The performance does plummet in the smaller versions, and even quantization may harm multilinguality disproportionally.

Something like deliberately distilling specific languages from the largest models could work well. Starting from scratch with a "legal" dataset will most likely fail as you say.

Silo AI (co-lead of this model) already tried Finnish and Scandinavian/Nordic models with the from-scratch strategy, and the results are not too encouraging.

https://huggingface.co/LumiOpen

Re: Open Euro LLM: Open LLMs for Transparent AI in Europe

#7

> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep r…

>As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be.

Train an LLM with text books and other legal books, you do not need to train it on pop culture to make it intelligent.

For face generations you might need to be more creative, you should not need milions of images stolen from social media to train your model.

But makes sense that tech giants do not want to share their data set and be transparent about stuff.

Re: Open Euro LLM: Open LLMs for Transparent AI in Europe

#9

> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep r…

I've been using Mistral past week due to changes in geopolitics, and Mistral works absolutely great in English. I haven't bothered in my native language yet, but in English it worked great. Better than my first experience with ChatGPT (GPT 3.5), actually.

Re: Open Euro LLM: Open LLMs for Transparent AI in Europe

#10

> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep r…

>As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. Train an LLM with text books and other legal books, you do not need to train it on pop culture to make it intelligent. For face generations you might need to be more creative, you should not need milions of images stolen from social media to train your model. But makes sense that tech giant…

> Train an LLM with text books and other legal books

Without licenses to the books, they are just as illegal (and maybe even moreso) than web content.

Post reply on HN