Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

41–50 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#41
post #17
post #13

As a Norwegian this sounds like a mistake. Who will use this LLM? Where? For what? The underlying data could be made more easily searchable and digestible for agents in general if the goal is better knowledge of Norwegian culture.

Exactly, if there's one thing transformers are good at it's translation. One I've found particularly nice: any question ChatGPT can answer in English it can answer in French. I'm assuming Norwegian too. So there's no point.

Yes transformers are great at translation as that is their purpose.

LLMs are not great at preserving cultural uniqueness and diversity. Take how “delve” has reentered the lexicon because the human assessors for pre training dialect of English uses “delve” a lot.

There is a lot of benefits to training specifically for a unique culture with unique norms to preserve the culture as we increasingly rely on LLMs.

https://www.scientificamerican.com/article/chatgpt-is-changi...

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#42

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

> this meager hardware

> they wasting - and why?

i18n language models are not area something frontier labs are focusing ton of resources on? ( certainly not in Norwegian)

The corpus of content in Norwegian - may not require very large clusters, or even if it does, this is best that the library could do, it would be certainly more than anyone else is investing in Norwegian models

SOTA models do not have the access to the quality of content that the national library does? The article mentions licensing with newspapers specifically, and the library has access to its own content archive.

English and Norwegian are not closely related language families, perhaps LoRA is not best approach?

I am curious if there is published research on how well localization works with LoRA depending on how far off the target language grammar/vocabulary is from English.

Projects like this typically have more than one objective and are not only building SOTA project, but is also to build/train foundational local talent , similar to universities launching satellites .

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#43
post #13

As a Norwegian this sounds like a mistake. Who will use this LLM? Where? For what? The underlying data could be made more easily searchable and digestible for agents in general if the goal is better knowledge of Norwegian culture.

I agree in principle.

That said, they are quite limited in what they are allowed to share of in-copyright works, and nb.no is a fantastic resource as it is (though you'll need a Norwegian IP address for too much of it - it's one of th main reasons I maintain a VPN) - if they are allowed to make it accessible there, it'd be great.

But they also have vast amounts of out-of-copyright data that I hope they'd make more easily accessible...

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#45
post #17
post #13

As a Norwegian this sounds like a mistake. Who will use this LLM? Where? For what? The underlying data could be made more easily searchable and digestible for agents in general if the goal is better knowledge of Norwegian culture.

Exactly, if there's one thing transformers are good at it's translation. One I've found particularly nice: any question ChatGPT can answer in English it can answer in French. I'm assuming Norwegian too. So there's no point.

Model can speak Lithuanian too, but with a Russian accent which is a big taboo for us.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#47
post #17

Earlier quoted context omitted.

Exactly, if there's one thing transformers are good at it's translation. One I've found particularly nice: any question ChatGPT can answer in English it can answer in French. I'm assuming Norwegian too. So there's no point.

They're only good at it because they were trained on massive amounts of English and French data.

Not really true.

Both Claude and ChatGPT can translate into minor dialects of Norwegian they will have seen very few works in because very few printed works exist in them.

E.g. I've tested both my local spoken dialect, which is rarely written, and a sociolect used by a 1970's Maoist group consiting of a few hundred people, where most of the printed material consists of novels from a couple of ex-members that became authors.

In the latter case, it claimed to not know, but was able to get a good match from just a description.

I also just had it ape Norwegian orthography from the 1910's by having it look up the rules and translate a text it had first translated from English to modern Norwegian, and it did just fine.

They will have seem some work in these dialects, but mostly it transfer really well to know related languages (English, Dutch, German, Swedish, Danish, roughly form a continuum from least in common to most in common with modern Norwegian; they all share vocabulary and significant parts of grammar with Norwegian), and then a relatively limited exposure to Norwegian itself is sufficient to do fairly well.

They're also really good at "style transfer" of text in the form of tweaking orthography, word order, and minor grammar changes from descriptions and examples.

(incidentally, the latter is one way of getting an LLM to sound a lot less like an LLM)

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#48
post #39
post #24

This can’t be right. 2 PB of flash is like $200k. It’s within reach of many individuals. Then again I guess you don’t need that much storage so maybe it is.

Your numbers are a little off but the point remains- 2PB is nothing, not newsworthy imo. What’s special about this?

What's special about it is not the flash but training an LLM based on the content, much of which is still in copyright and which the library has restrictions on how they are allowed to use (irrespective of the legal position of training on it) and which required an agreement with the copyright holders.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#49

384 core cpu cluster? 2 petabytes? Dell just launched a 2U that fits almost 10 petabytes in it. It's probably not 384 core capable but that is very doable right now, Epyc chips are 192 cores each! https://www.techradar.com/pro/dell-launches-record-shatterin...

That's the in-house preprocessing hardware, not what they're training on.

Yes!

It's still a weird article, to highlight a "big" storage appliance. Having all that NVMe local feels like it would be much much much much faster.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#50

That's about 350MB per capita. Humans can produce 2-6kb per hour. That's 13 years of non-stop typing. Wonder where it all comes from. I guess it's websites that aren't compressed / extracted.

It's a legal deposit library, same as e.g. Library of Congress. Which means almost every published book, magazine, and newspaper and many other works published in Norway, as well as large collections of Norwegian works published abroad (such as thousands of Norwegian-language newspapers published by the Norwegian immigrant communities in the US) for many decades and a large proportion of the same from the last 200+ years are stored there.

They do also crawl websites (or at least did) in the .no tld.

Post reply on HN