Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

51–60 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#51
post #16

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

That's what they have access to right now. I am sure that will change in the future as the project progresses. What do you suggest, that they stop and wait until they have the right HW?

Also, it's Norway...

"Norway's sovereign wealth fund, officially known as the Government Pension Fund Global, is the world's largest sovereign wealth fund with assets exceeding \(\$2\) trillion. Established in 1990 and managed by Norges Bank Investment Management, it was created to channel surplus petroleum revenues into long-term global investments to benefit future generations."

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#52

I'm a Norwegian, and I use the national library almost every day for searching through texts. They have truly one of the best working user interfaces (and functionality) for searching through the massive amounts of text.

It's really fantastic. I just wished there were fewer restrictions on the content that is accessible.

(a lot is only accessible from Norwegian IP addresses, so it's one of the main reasons I maintain a VPN as I'm Norwegian but live in the UK; a second set is only available from the IP addresses of libraries or research institutions - still huge amounts that are generally available, though)

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#53
post #8

>As Husnes put it; Norway is a small country solving a problem every non-English-speaking nation will face: how do you build AI that reflects your language, your culture and your history? AI needs custodians, not just builders. I'm afraid the answer is, mostly you don't. Such a thing requires strong political will that, at least in my environment, seems basically impossible to align. The costs are prohibitive, but be…

I guess it's subject to debate whether the cost indeed is prohibitive in the case of Norway. They are a small but extremely wealthy country - after all, they currently hold the equivalent of 1,5% of all the listed companies globally through the investments of their sovereign wealth fund.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#54
How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language."

I thought all big players already train on basically everything remotely available to them no matter the language or quality, so his take sounds like an opinion formed in the early days of generally available LLMs.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#55
post #42

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

> this meager hardware > they wasting - and why? i18n language models are not area something frontier labs are focusing ton of resources on? ( certainly not in Norwegian) The corpus of content in Norwegian - may not require very large clusters, or even if it does, this is best that the library could do, it would be certainly more than anyone else is investing in Norwegian models SOTA models do not have the access to…

> English and Norwegian are not closely related language families, perhaps LoRA is not best approach?

Yes, they are. English is a West Germanic language. Norwegian is a North Germanic language. The French vocabulary in English obscures it a bit, but the two languages have similar grammar and the vocabulary has a huge number of close cognates.

E.g. day -> dag, ship -> skip, apple -> eple, cow -> ku (which makes more sense when you pronounce them correctly out loud), bairn (child; mostly Scotland and Northern England) -> barn, hop -> hopp, yule -> jul just to give a random selection of English Germanic words.

But more than that, the frontier models both a) knows Norwegian quite well, b) certainly knowns German and Dutch well, and there's a continuum of language transfer around the North sea especially when accounting for sounds rather than modern orthography, e.g. to take a couple of examples from above: ship -> schip -> Schiff -> skib -> skip; day -> dag -> Tag -> dag). The "jump" to Dutch already weeds out most of the French. A lot of modern Norwegian orthography comes from Danish, which again shares more than modern Norwegian does with German.

Knowing any of these helps a lot with learning Norwegian and vice versa. E.g. I'm Norwegian, I've never learnt Dutch, but I have learnt English and German, and I can read Dutch fairly well from that alone.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#56

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

> Which begs the question, whose money are they wasting - and why?

Norway is better run as a country than 99% of the countries on the planet, including the one that invented current LLM tech, so I'd give them the benefit of the doubt.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#57

How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…

If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#59
post #55
post #42

Earlier quoted context omitted.

> this meager hardware > they wasting - and why? i18n language models are not area something frontier labs are focusing ton of resources on? ( certainly not in Norwegian) The corpus of content in Norwegian - may not require very large clusters, or even if it does, this is best that the library could do, it would be certainly more than anyone else is investing in Norwegian models SOTA models do not have the access to…

> English and Norwegian are not closely related language families, perhaps LoRA is not best approach? Yes, they are. English is a West Germanic language. Norwegian is a North Germanic language. The French vocabulary in English obscures it a bit, but the two languages have similar grammar and the vocabulary has a huge number of close cognates. E.g. day -> dag, ship -> skip, apple -> eple, cow -> ku (which makes more s…

This makes me deeply curious about how LLMs understand language. Do LLMs relate cognates more than words that are dissimilar in different languages? I wonder if that plays some role in the effectiveness of tokenization.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#60
post #40
post #32

I wonder if instead (or in parallel), Norway should build a set of training data and share it (for free) with all the model builders. Seems like making the frontier models know Norwegian and their culture is a better (or additional!) way to reach the end they are going for here.

The frontier models know Norwegian just fine. They can also adapt to Norwegian dialects, and even ape old Norwegian fairly well. E.g. I had Claude describe the novel "De knyttede næver" from 1911 in Norwegian orthography ca. 1911, as it's a novel I've read, and it does a good job. What it lacks is an understanding of Norwegian literature, culture and history. It had to look up "De knyttede næver", which was one of th…

Odd, I'd imagine Wikisource (in many/all languages) would be part of training data for all LLMs with SOTA ambition?

https://no.wikisource.org/wiki/De_knyttede_n%C3%A6ver

Post reply on HN