Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

61–70 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#61

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

"Training a sovereign LLM with this meager hardware"

Norway has a sovereign fund worth O[MS|Apple|etc] except it is largely in readies and not pixie dust.

Whilst the UK frittered away North Sea oil profits, Norge squirreled them away instead.

So, if the grand dream of LLMs and AI does actually come to some sort of fruition and not simply another case of the Emperor's New Clothes combined with some lovely tulips and a dotcom boom and bust, then Norge can simply stuff shit loads of cash into buying whatever they need. Cash is king after all.

The beast they have described here is just a library system. I think I'd like my country's (UK) library system to have resources like that.

I don't think you are asking the right question: When you say "meager", I see "rather impressive PoC from a well resourced organisation"

You say tomato ...

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#62

How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…

If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

absolutely. somebody online was wanting an LLM with Georgian language support, and that's exactly what i suggested: start digitizing Georgian text.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#63
post #55

Earlier quoted context omitted.

> English and Norwegian are not closely related language families, perhaps LoRA is not best approach? Yes, they are. English is a West Germanic language. Norwegian is a North Germanic language. The French vocabulary in English obscures it a bit, but the two languages have similar grammar and the vocabulary has a huge number of close cognates. E.g. day -> dag, ship -> skip, apple -> eple, cow -> ku (which makes more s…

This makes me deeply curious about how LLMs understand language. Do LLMs relate cognates more than words that are dissimilar in different languages? I wonder if that plays some role in the effectiveness of tokenization.

I have no idea if the similar spelling will somehow help - I used that mostly because it's a simple way if illustrating the close relationship, but I suspect you'd find that the meanings of closely related words are likely to more directly overlap.

The grammar is perhaps more likely to help. Similar word order etc. Even weirdness like German - my only top grade on a German essay in school was one where I on purpose ignored what I thought I knew about German and tried to evoke "old fashioned" Norwegian. The result was guessing at a bunch of grammatical structures that I didn't know if was valid German. Turned out I was right about most of it - century old Norwegian was far closer to century old Danish, was a lot closer to valid German, and enough so to impress my teacher enough to overlook a number of orthographic mistakes.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#64
Even entire governments are captured by a mild LLM psychosis. Which is sad in the case of Norway. I lived in Norway for two years and always found their government to be highly rational, this is not a rational use of public funds (but I suppose they have plenty of capital).

Western society is completely captured by this form of psychosis and its going to bite us in the a* very soon.

I firmly believe all the Boomer leaders throughout the world are being sold a bag of lies by technocrats that "AI", specifically LLMs, are going to cure disease and death and therefor they are willing to handover all control to the technocrats. Fckin croakers at it again.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#65

How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…

If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

Yeah, was about to comment that too, instead of training a new model and new weights exclusively for Norwegian (and expecting/wanting every other small/medium-sized country to do the same) which seems infinity harder, they could have made high quality transcriptions and translations of the stories currently described only in Norwegian into English, and making it all public. I guess there still would be a worry that it'd be counted as "less important" compared to other history, news and culture about other countries.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#66
post #60
post #40

Earlier quoted context omitted.

The frontier models know Norwegian just fine. They can also adapt to Norwegian dialects, and even ape old Norwegian fairly well. E.g. I had Claude describe the novel "De knyttede næver" from 1911 in Norwegian orthography ca. 1911, as it's a novel I've read, and it does a good job. What it lacks is an understanding of Norwegian literature, culture and history. It had to look up "De knyttede næver", which was one of th…

Odd, I'd imagine Wikisource (in many/all languages) would be part of training data for all LLMs with SOTA ambition? https://no.wikisource.org/wiki/De_knyttede_n%C3%A6ver

You'd think so. It seems like there are a lot of odd gaps like that.

I also have a favourite English language PhD thesis I ask every new model about that they still struggle to find even though there's a Wikipedia article about it that links a blog post I wrote about it.

Anyone who thinks they've exhausted even publicly crawlable resources should ask them about some obscure stuff.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#67
post #8

>As Husnes put it; Norway is a small country solving a problem every non-English-speaking nation will face: how do you build AI that reflects your language, your culture and your history? AI needs custodians, not just builders. I'm afraid the answer is, mostly you don't. Such a thing requires strong political will that, at least in my environment, seems basically impossible to align. The costs are prohibitive, but be…

I'm sure if Norway approached the American labs with goal of making a curated datasets for training, they would absolutely get in the training door, and those models would likely run circles around anything that could be domestically done. That being said though, I can feel you cringing through the screen.

>That being said though, I can feel you cringing through the screen.

Then I failed to express myself in writing. I'm definitely a fan of this kind of initiative and am not happy with the type of viability I think they have.

I might very well be projecting a whole lot of local dynamics of national identity, politics and culture though.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#68

I'm a Norwegian, and I use the national library almost every day for searching through texts. They have truly one of the best working user interfaces (and functionality) for searching through the massive amounts of text.

The lack of a universal search engine is very frustrating. Why can't I search within TV subtitles?

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#69

How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…

If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

What incentives does OpenAI have to make sure the AI actually works well with Norwegian beyond capturing a (small) Norwegian market? What incentives do they have to take Norwegian values into consideration, or to preserve Norwegian culture into the future? The matter is also a question of national sovereignty, so to simply release the data and nicely ask foreign companies to solve the problem for you, would be a fool's move

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#70
post #61

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

"Training a sovereign LLM with this meager hardware" Norway has a sovereign fund worth O[MS|Apple|etc] except it is largely in readies and not pixie dust. Whilst the UK frittered away North Sea oil profits, Norge squirreled them away instead. So, if the grand dream of LLMs and AI does actually come to some sort of fruition and not simply another case of the Emperor's New Clothes combined with some lovely tulips and a…

The reason they have the largest sovereign wealth fund (aside from getting it right in the 80s, unlike the UK), is that there is quite a bit of regulation around where and how the money is invested.

It is run to maximise growth for example, so even though Norway is way ahead with electric car usage and infrastructure (presumably because they have a climate likely to be most affected by global warming/heating) their fund still invests in fossil fuels as they are a profit/growth opportunity.

Anyway, i don't think it's as easy as "simply stuff shit loads of cash into buying whatever they need". I believe there would be a serious political discussion needed for that to happen.

Post reply on HN