Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

131–140 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#131
post #86

> Marius Husnes, the Head of IT Platform at the library (Nasjonlbiblioteket) discussed the project at Huawei’s ID Forum 2026 in Paris, saying that no commercial LLM provider was developing a local (Norwegian) language LLM. He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about…

He’s right though, although it’s not entirely about the training corpus. It’s about the tokenizer that tokenizes substrings more efficiently based on a necessary bias towards a target language. English oriented LLMs are more powerful for English than other languages because the token space is more parsimonious in English language. Try any online Anthropic tokenizer that calls their api with common English words (typi…

Did you even try to verify your claims. I tested it on few translations on wikipedia articles using [1] and it takes 15-20% more tokens for Norwegian.

English performs the best because there is more data in English and high quality sources are either only in English or there is a good translation in English.

[1]: https://platform.openai.com/tokenizer

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#133
post #52

I'm a Norwegian, and I use the national library almost every day for searching through texts. They have truly one of the best working user interfaces (and functionality) for searching through the massive amounts of text.

It's really fantastic. I just wished there were fewer restrictions on the content that is accessible. (a lot is only accessible from Norwegian IP addresses, so it's one of the main reasons I maintain a VPN as I'm Norwegian but live in the UK; a second set is only available from the IP addresses of libraries or research institutions - still huge amounts that are generally available, though)

Silly question but can a non-Norwegian also access it? Willing to pick up some Norwegian along the way ;-)

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#134

Earlier quoted context omitted.

The issue is that French, Italian, African, Japanese people shouldn't have the inconvenience of instructing the LLM tool to get the basic facts about their own culture. They should use an LLM that has already been trained like that by default. Nobody has obligation to use a tool that thinks it is talking to an American. If I go to Google for example I want to get facts about my own country in my own language.

Wouldn't those people be asking the questions in their own language in the first place? The model will reply in the language you use. This thread is about people asking for information about a language that is not the one they are messaging the LLM in

Even if the model will reply in my language, I often notice it searching in english. Or thinking in english. There's always something lost in translation. Sometimes it's just minor nuances. Other times it mangles the legal facts with those of other countries.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#135

Earlier quoted context omitted.

The issue is that French, Italian, African, Japanese people shouldn't have the inconvenience of instructing the LLM tool to get the basic facts about their own culture. They should use an LLM that has already been trained like that by default. Nobody has obligation to use a tool that thinks it is talking to an American. If I go to Google for example I want to get facts about my own country in my own language.

> Nobody has obligation to use a tool that thinks it is talking to an American. Then add top-level instructions saying what country you're from, what country you live in now, and which language you speak. This isn't that hard.

None of that even addresses the problem described, because none of the languages you mentioned would be French in the described example.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#136

Earlier quoted context omitted.

It's also a bit funny because Norway definitely has enough money to hire a team of Anthropic's best to go out there and train them a model that does whatever they want. They probably have enough money to fund their own Anthropic competitor.

>They probably have enough money to fund their own Anthropic competitor. Which is bizarre to me Norway doesn't have a booming tech sector with all hat wealth fund acting as the biggest VC. They instead use their wealth fund to invest in US's tech sector. Baffling.

There's only so much you can do with 5 million people. Especially in a field where network effects amd scale matter a lot.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#137

Earlier quoted context omitted.

This model is going to start miles behind the frontier and the gap will only grow.

Why would the gap grow? There is no more training data to acquire, frontier model are training on the entire internet. Everything from now on is just fine-tuning.

Your statement assumes training data is the only thing that matters for the big players, while not considering it limiting for the small Norwegian model. That’s a fallacy.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#138
post #69

Earlier quoted context omitted.

What incentives does OpenAI have to make sure the AI actually works well with Norwegian beyond capturing a (small) Norwegian market? What incentives do they have to take Norwegian values into consideration, or to preserve Norwegian culture into the future? The matter is also a question of national sovereignty, so to simply release the data and nicely ask foreign companies to solve the problem for you, would be a fool…

It's also a bit funny because Norway definitely has enough money to hire a team of Anthropic's best to go out there and train them a model that does whatever they want. They probably have enough money to fund their own Anthropic competitor.

I highly doubt that hiring people who don't even speak the language would result in a better model for Norwegian. If anything, they could pay Anthropic for some tips and tricks for training. But that does not seem necessary as Deepseek & co detail everything for free

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#139

Earlier quoted context omitted.

>They probably have enough money to fund their own Anthropic competitor. Which is bizarre to me Norway doesn't have a booming tech sector with all hat wealth fund acting as the biggest VC. They instead use their wealth fund to invest in US's tech sector. Baffling.

There's only so much you can do with 5 million people. Especially in a field where network effects amd scale matter a lot.

Finland has same population as Norway, has way less money, but has 3x the scaleups. Even bigger difference with vs Netherlands.

Even Norway themselves admit they're the underperformers of the Nordics. https://skywlkr.no/wp-content/uploads/2019/10/TechScaleupNor...

So blaming population is a cheap excuse that doesn't hold water. Especially that you can always import the skilled people you lack, when you have virtually unlimited money and some of the highest standards of living in the world.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#140
post #26
post #17

Earlier quoted context omitted.

Exactly, if there's one thing transformers are good at it's translation. One I've found particularly nice: any question ChatGPT can answer in English it can answer in French. I'm assuming Norwegian too. So there's no point.

The point is that norway willl have its own LLM. And will not have dependencies to another state or private company. The goal is not to be the best model. But to have a model that include more Norwegian data then other LLM and that it's not screwed against other sources.

But what does that give you? If the model is far less capable? What will it do for you with that Norwegian data, that a better model could not do with better search or context?
Post reply on HN