Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

121–130 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#122
post #17
post #13

As a Norwegian this sounds like a mistake. Who will use this LLM? Where? For what? The underlying data could be made more easily searchable and digestible for agents in general if the goal is better knowledge of Norwegian culture.

Exactly, if there's one thing transformers are good at it's translation. One I've found particularly nice: any question ChatGPT can answer in English it can answer in French. I'm assuming Norwegian too. So there's no point.

Definitively not the case in distilled models.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#123
post #47

Earlier quoted context omitted.

Not really true. Both Claude and ChatGPT can translate into minor dialects of Norwegian they will have seen very few works in because very few printed works exist in them. E.g. I've tested both my local spoken dialect, which is rarely written, and a sociolect used by a 1970's Maoist group consiting of a few hundred people, where most of the printed material consists of novels from a couple of ex-members that became a…

This is all true, but I assumed the original posters were talking about cultural knowledge, not linguistic correspondences. To do translation well you still need cultural knowledge. (E.g. the particular modes of specific kinds of legalese, or slang and the nuances of social class, etc)

I think it's not that this knowledge isn't present in the model somewhere, but probably more that it gets killed by instruction tuning for US corporate values.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#124
post #69

Earlier quoted context omitted.

What incentives does OpenAI have to make sure the AI actually works well with Norwegian beyond capturing a (small) Norwegian market? What incentives do they have to take Norwegian values into consideration, or to preserve Norwegian culture into the future? The matter is also a question of national sovereignty, so to simply release the data and nicely ask foreign companies to solve the problem for you, would be a fool…

It's also a bit funny because Norway definitely has enough money to hire a team of Anthropic's best to go out there and train them a model that does whatever they want. They probably have enough money to fund their own Anthropic competitor.

>They probably have enough money to fund their own Anthropic competitor.

Which is bizarre to me Norway doesn't have a booming tech sector with all hat wealth fund acting as the biggest VC.

They instead use their wealth fund to invest in US's tech sector. Baffling.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#125

Earlier quoted context omitted.

I task GPT/Claude with researching stuff that pertains to very specific cultural or legal aspects in French politics, on a daily basis. Even though French is a way more common language globally than Norwegian, these models still haven't figured out that, no matter the language I myself speak to them (German or English depending on my mood) their web searches need to be done in French to return reasonable results. I h…

Aren’t you already using English in the LLM convo? Telling the model to use French for research or to find resources in French seems like a reasonable step. If you’re doing this on a daily basis, then you should have an AGENTS.md that accumulates directional instructions like this. This is how you use the tool correctly. There’s this weird pattern I’ve noticed where people expect LLMs to require zero effort or profic…

> Aren’t you already using English in the LLM convo? Telling the model to use French for research or to find resources in French seems like a reasonable step.

Most ordinary people will just use their native language and they have no way of knowing that the model always reasons in English and therefore is strongly biased toward using English search terms. So they don't know they have to remind the model to search in their local language.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#127
post #86

> Marius Husnes, the Head of IT Platform at the library (Nasjonlbiblioteket) discussed the project at Huawei’s ID Forum 2026 in Paris, saying that no commercial LLM provider was developing a local (Norwegian) language LLM. He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about…

[dead]

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#129

How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…

Quite true ? English is ludicrously over abundant in training when compared to any language.

And that's probably necessary if you want a competent model. There simply isn't much norwegian literature on let's say banana farming.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#130
post #36

Earlier quoted context omitted.

More like $1M at current prices at this scale / level of performance. If you go with HDD arrays probably $50k

Boy pricing is pretty nuts these days. I have half a petabyte in Seagate enterprise drives myself and I didn’t pay anything close to that to acquire it. Such a pity about the flash storage. 2 years ago we built 200 TiB or something of flash using Samsung PM1633 or something and it was a fraction of the cost per gigabyte that $1m would imply.

We're in the boom phase of the cycle. The bust on these chips always comes.
Post reply on HN