> The models will be developed within Europe's robust regulatory framework, ensuring alignment with European values while maintaining technological excellence As a European, that's practically an oxymoron. The more one limits oneself to legally clean data, the worse the models will be. I hate to be pessimistic from the get go, but it doesn't sound like anything useful will be produced by this and we'll have to keep r…
What do you mean by relying on Google? Llama 3.1 and DeepSeek v3/R1 largest models are rather good at even a niche language like Finnish. The performance does plummet in the smaller versions, and even quantization may harm multilinguality disproportionally. Something like deliberately distilling specific languages from the largest models could work well. Starting from scratch with a "legal" dataset will most likely f…
Something like 4o is so perfect in most languages that one could just make an infinite dataset from it and be done with it. I'm not sure how OAI managed it tbh.