Live data from Hacker News

AMÁLIA and the future of European Portuguese LLMs

duarteocarmo.com

21–30 of 90 posts

Re: AMÁLIA and the future of European Portuguese LLMs

#21
post #17

Earlier quoted context omitted.

Right, and my point is that if you use 80% Brazilian Portuguese during base model training + 20% European Portuguese as post-training, you pretty much get exactly that, except with a ton more of available training data.

What's your evidence for that? And if the first 80% doesn't bias the language after post-training (which I think is what you're claiming) why not go for English or a mixture of languages, which is essentially what they did by starting with EuroLLM?

Evidence? Not so much, I didn't realize I was defending a PhD thesis here.

I speak Spanish, and have talked with people who only speak Portuguese, either of the variants, and also talked with Portuguese people before how they see their language, comparing it with Brazilian Portuguese, and vice-versa. So basically based on vibes and experience.

> And if the first 80% doesn't bias the language after post-training (which I think is what you're claiming) why not go for English

I'm not sure how many languages you speak or encountered in the wild before, but some languages are VERY different from each other, some are a bit different and others are basically the same with some differences. Doing what I describe for languages that are similar is easier than languages that are very different, for what I hope are obvious reasons.

Re: AMÁLIA and the future of European Portuguese LLMs

#22

Earlier quoted context omitted.

I don't think so, Portugal the country might be small, with a small population, but there is ~250 million "Lusophones" (native Portuguese speakers), making it the fifth-most spoken native language in the world, I'd hardly call that small :) And before everyone screams; yes, European Portuguese is different from Brazilian Portuguese, but they're still both Portuguese and understand each other, so it's not like the tex…

The authors are pretty clearly trying to draw only from European Portuguese sources - I feel like there's a fairly widespread attitude here that the language is being overwhelmed by the sheer number of Brazilian speakers (which there is obviously at least some truth to). I don't necessarily personally feel like preserving European Portuguese in amber is a worthwhile goal (anymore than it is productive for Brits to be…

> I don't necessarily personally feel like preserving European Portuguese in amber is a worthwhile goal (anymore than it is productive for Brits to be prickly about the meteoric rise of US English).

That's easy to say when you're not on the other end of US defaultism.

Re: AMÁLIA and the future of European Portuguese LLMs

#23

It is definitely an interesting problem, because Portugal is a small enough country that the actual total corpus of available texts in (non-Brazilian) Portuguese is potentially problematic.

I don't think so, Portugal the country might be small, with a small population, but there is ~250 million "Lusophones" (native Portuguese speakers), making it the fifth-most spoken native language in the world, I'd hardly call that small :) And before everyone screams; yes, European Portuguese is different from Brazilian Portuguese, but they're still both Portuguese and understand each other, so it's not like the tex…

Right, but most of those speak brazilian portuguese. There's so much less european portuguese text that it becomes impossible for a model to not speak brazilian portuguese if not trained in a way that ignores brazilian sources

Re: AMÁLIA and the future of European Portuguese LLMs

#24
post #17

Earlier quoted context omitted.

What's your evidence for that? And if the first 80% doesn't bias the language after post-training (which I think is what you're claiming) why not go for English or a mixture of languages, which is essentially what they did by starting with EuroLLM?

Evidence? Not so much, I didn't realize I was defending a PhD thesis here. I speak Spanish, and have talked with people who only speak Portuguese, either of the variants, and also talked with Portuguese people before how they see their language, comparing it with Brazilian Portuguese, and vice-versa. So basically based on vibes and experience. > And if the first 80% doesn't bias the language after post-training (whic…

> I'm not sure how many languages you speak or encountered in the wild before, but some languages are VERY different from each other, some are a bit different and others are basically the same with some differences.

I'm a dual citizen of Portugal and Brazil and I live in the US now, so that's my linguistic background. (Also studied bits of French, Russian, Latin and Greek.)

> Doing what I describe for languages that are similar is easier than languages that are very different, for what I hope are obvious reasons.

Not only are your reasons not obvious, your conclusion is actually wrong.

If the goal is to create an LLM with minimal Brazilian Portuguese bias (which was one of their main goals), it might actually make more sense to train it in any other language BUT Brazilian Portuguese (say, English), then fine-tune it for European Portuguese.

LLM's have shown to be very good at generalizing across languages (the transformer architecture literally comes from work on translators IIRC).

Re: AMÁLIA and the future of European Portuguese LLMs

#26

It is definitely an interesting problem, because Portugal is a small enough country that the actual total corpus of available texts in (non-Brazilian) Portuguese is potentially problematic.

European Portuguese is the 13th most populous language in Europe. Not that small, there are many other European languages in use that are much smaller.

https://en.wikipedia.org/wiki/List_of_languages_by_number_of...

Re: AMÁLIA and the future of European Portuguese LLMs

#27
post #26

It is definitely an interesting problem, because Portugal is a small enough country that the actual total corpus of available texts in (non-Brazilian) Portuguese is potentially problematic.

European Portuguese is the 13th most populous language in Europe. Not that small, there are many other European languages in use that are much smaller. https://en.wikipedia.org/wiki/List_of_languages_by_number_of...

> European Portuguese is the 13th most populous language in Europe

that's not impressive

Re: AMÁLIA and the future of European Portuguese LLMs

#28
post #26

Earlier quoted context omitted.

European Portuguese is the 13th most populous language in Europe. Not that small, there are many other European languages in use that are much smaller. https://en.wikipedia.org/wiki/List_of_languages_by_number_of...

> European Portuguese is the 13th most populous language in Europe that's not impressive

Hello from 23rd

Re: AMÁLIA and the future of European Portuguese LLMs

#29
post #7

I'm not sure the direction should be to finetune a small local model for each country or language. These models are already not particularly great at information retrieval, so I doubt anyone would use them for questions like the author suggests (ie who was the president between X and Y). Similarly, they are a little too lightweight to be used for translations too. If the budget is indeed so modest (5.5 million euros!…

I agree, the research is complex enough as is without having to worry about splitting it babel-like into multiple languages.

Re: AMÁLIA and the future of European Portuguese LLMs

#30
post #7

I'm not sure the direction should be to finetune a small local model for each country or language. These models are already not particularly great at information retrieval, so I doubt anyone would use them for questions like the author suggests (ie who was the president between X and Y). Similarly, they are a little too lightweight to be used for translations too. If the budget is indeed so modest (5.5 million euros!…

Yeah I think India is going the better route with Sarvam which is trained from scratch and still relatively cheap.
Post reply on HN