Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

31–40 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#31

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

The largest problem is available training data actually.

They have already done experiments with dittrent sub 10b models with both fine-tuning and fully from scratch. And last I check the fully from scratch captured the language in a better way.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#32
I wonder if instead (or in parallel), Norway should build a set of training data and share it (for free) with all the model builders.

Seems like making the frontier models know Norwegian and their culture is a better (or additional!) way to reach the end they are going for here.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#33
post #8

>As Husnes put it; Norway is a small country solving a problem every non-English-speaking nation will face: how do you build AI that reflects your language, your culture and your history? AI needs custodians, not just builders. I'm afraid the answer is, mostly you don't. Such a thing requires strong political will that, at least in my environment, seems basically impossible to align. The costs are prohibitive, but be…

I'm sure if Norway approached the American labs with goal of making a curated datasets for training, they would absolutely get in the training door, and those models would likely run circles around anything that could be domestically done.

That being said though, I can feel you cringing through the screen.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#34
post #20
post #2

2 PB? They will not come close to training in on that amount. Maybe years from now.

Think they will not train on the dull 2TB but use that as the data lake to start and then apply a more targeted approach.

if you read the article 2pb is available as flash storage in the data pipeline, used to dedupe, clean, normalize, etc, for training from 60pb of raw data.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#35

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

They successfully have made PoC finetunes before, so the next step is training fully fledged LLMs.

I don’t think they aim to anything worthwhile. The finetunes were incredibly broken. I’m guessing it’s more about having the method to do it. I’m not convinced it’s super useful but I’m not one to decide who gets to do what with the research funds.

One finetune I tried did make fun of humans expressing their feelings in the chat. Often.

One other finetune did hallucinate that it was a doctor and my baby had terrible diseases, every time I just wrote "hei" (with a generic neutral system prompt that likely triggered this behaviour though).

I think Olivia is big enough for what it’s used for. In my opinion it’s better to stay up to date and not waste too much money on hardware at the moment.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#36
post #24

This can’t be right. 2 PB of flash is like $200k. It’s within reach of many individuals. Then again I guess you don’t need that much storage so maybe it is.

More like $1M at current prices at this scale / level of performance. If you go with HDD arrays probably $50k

Boy pricing is pretty nuts these days. I have half a petabyte in Seagate enterprise drives myself and I didn’t pay anything close to that to acquire it. Such a pity about the flash storage. 2 years ago we built 200 TiB or something of flash using Samsung PM1633 or something and it was a fraction of the cost per gigabyte that $1m would imply.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#37

384 core cpu cluster? 2 petabytes? Dell just launched a 2U that fits almost 10 petabytes in it. It's probably not 384 core capable but that is very doable right now, Epyc chips are 192 cores each! https://www.techradar.com/pro/dell-launches-record-shatterin...

That's the in-house preprocessing hardware, not what they're training on.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#39
post #24

This can’t be right. 2 PB of flash is like $200k. It’s within reach of many individuals. Then again I guess you don’t need that much storage so maybe it is.

Your numbers are a little off but the point remains- 2PB is nothing, not newsworthy imo. What’s special about this?

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#40
post #32

I wonder if instead (or in parallel), Norway should build a set of training data and share it (for free) with all the model builders. Seems like making the frontier models know Norwegian and their culture is a better (or additional!) way to reach the end they are going for here.

The frontier models know Norwegian just fine. They can also adapt to Norwegian dialects, and even ape old Norwegian fairly well.

E.g. I had Claude describe the novel "De knyttede næver" from 1911 in Norwegian orthography ca. 1911, as it's a novel I've read, and it does a good job.

What it lacks is an understanding of Norwegian literature, culture and history. It had to look up "De knyttede næver", which was one of the best-selling Norwegian novels around the time it was published before I'd get anything out of it (ChatGPT does better; in thinking mode in particular it gives a detailed summary).

While not exactly well known today, the author was a prominent newspaper journalist for decades, and the novel series is well enough known that e.g. there's a Norwegian singer that took his stage name after the protagonist, and it was covered in Norwegian papers and books for decades (partly because of controversy over the authors political views and how they coloured his novels), so it does feel like a reasonable test that reveals a quite significant knowledge gap.

I do agree with you that it'd be better if the data set from the national library was made more accessible, though it seems a major addition here is that they have a deal to train on copyrighted data locked away in their archives that they have limitations on the use of.

But even just making the out of copyright data in their collections would be a great start.

Post reply on HN