Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

91–100 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#91
post #66
post #60

Earlier quoted context omitted.

Odd, I'd imagine Wikisource (in many/all languages) would be part of training data for all LLMs with SOTA ambition? https://no.wikisource.org/wiki/De_knyttede_n%C3%A6ver

You'd think so. It seems like there are a lot of odd gaps like that. I also have a favourite English language PhD thesis I ask every new model about that they still struggle to find even though there's a Wikipedia article about it that links a blog post I wrote about it. Anyone who thinks they've exhausted even publicly crawlable resources should ask them about some obscure stuff.

the models don't retain their full training data set

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#92

Earlier quoted context omitted.

If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

Yeah, was about to comment that too, instead of training a new model and new weights exclusively for Norwegian (and expecting/wanting every other small/medium-sized country to do the same) which seems infinity harder, they could have made high quality transcriptions and translations of the stories currently described only in Norwegian into English, and making it all public. I guess there still would be a worry that i…

> high quality transcriptions and translations of the stories currently described only in Norwegian into English

You make it sound like an easier task than training an LLM. I'd argue it's not obvious, and would assume the contrary.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#93
post #13

As a Norwegian this sounds like a mistake. Who will use this LLM? Where? For what? The underlying data could be made more easily searchable and digestible for agents in general if the goal is better knowledge of Norwegian culture.

Hard disagree. This is the first step not the last and proves to other countries that this can be done.

This model is going to start miles behind the frontier and the gap will only grow.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#94

Earlier quoted context omitted.

Aren’t you already using English in the LLM convo? Telling the model to use French for research or to find resources in French seems like a reasonable step. If you’re doing this on a daily basis, then you should have an AGENTS.md that accumulates directional instructions like this. This is how you use the tool correctly. There’s this weird pattern I’ve noticed where people expect LLMs to require zero effort or profic…

The issue is that French, Italian, African, Japanese people shouldn't have the inconvenience of instructing the LLM tool to get the basic facts about their own culture. They should use an LLM that has already been trained like that by default. Nobody has obligation to use a tool that thinks it is talking to an American. If I go to Google for example I want to get facts about my own country in my own language.

> Nobody has obligation to use a tool that thinks it is talking to an American.

Then add top-level instructions saying what country you're from, what country you live in now, and which language you speak. This isn't that hard.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#95
post #86

> Marius Husnes, the Head of IT Platform at the library (Nasjonlbiblioteket) discussed the project at Huawei’s ID Forum 2026 in Paris, saying that no commercial LLM provider was developing a local (Norwegian) language LLM. He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about…

You're making the mistake of thinking whether he knows what he is talking about matters. He is brewing a potion. It's ingredients are a trendy term, a vaguely spooky threat and a clear, overly simplistic solution that of course he will graciously assume control of, for the good of the motherland.

This potion is potent and you'd think it would stop working from frequent misuse but you'd be wrong!

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#96

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

That's enough resources to build on something like the Olmo 3 recipe but with a mix prioritizing their own data and post-training for their own tasks. If they build their own embedding model, index everything in the library, and train their model to query that data while answering historical, cultural, legal, and strategic questions from their perspective... Pretty interesting and likely useful. They won't beat Anthropic at dumping out React code but also there's no real reason to duplicate that.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#97

Earlier quoted context omitted.

Current-best models are pretty fluent at major languages and cultures, so it's untrue at least for the "any" qualifier. Performance is barely affected or might be even better sometimes. However English patterns can subtly leak into native patterns of other languages. It's obviously very different for low-resource languages, but to improve them you need more data, not a new model.

>Current-best models are pretty fluent at major languages and cultures strong disagree on that one. As a German interacting with ChatGPT, even in German it gives me the feeling of talking to the Pluribus people, which reminds me of an anecdote of Walmart failing in Germany because people were freaked out by the constantly upbeat, smiling employees. Understanding a culture is a very different task than translating the…

I'm Finnish and dear god I hate the default overtly friendly tones of LLMs. Always the first thing to tune in system prompt.

You're a machine, stop anthropomorphizing yourself and pretending to be my best friend, and just give me the damn answer and nothing else. :D

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#98

How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…

As the article explains, Norway's National Library has a database of practically everything published and broadcast in Norwegian going back many decades. From the way the dataset described in the article, it does not sound like OpenAI et al. would have easy access to it in its entirety.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#99
post #86

> Marius Husnes, the Head of IT Platform at the library (Nasjonlbiblioteket) discussed the project at Huawei’s ID Forum 2026 in Paris, saying that no commercial LLM provider was developing a local (Norwegian) language LLM. He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about…

You're making the mistake of thinking whether he knows what he is talking about matters. He is brewing a potion. It's ingredients are a trendy term, a vaguely spooky threat and a clear, overly simplistic solution that of course he will graciously assume control of, for the good of the motherland. This potion is potent and you'd think it would stop working from frequent misuse but you'd be wrong!

He won't have control over it.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#100
post #19

Earlier quoted context omitted.

They made the cultural case, you have no idea how strong this is in places like quebec, nordics, france, russia etc

Can confirm that. Norway may have a small population, but if you live there you'll think it's truly the center of the world (aside from the US. Norwegians love America)

Love America? Yes, we did.
Post reply on HN