Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

201–210 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#201
post #101

Earlier quoted context omitted.

When I am chatting with ChatGPT - it is fairly obvious that it is American - its native language, its style, its attitude is American - even if we chat in Danish. Just as we cannot rely on Netflix and HBO to produce Scandinavian TV-shows even though they might do at the moment, we need to make our own stuff in this area too. And over time, the technology to do this will become cheap and readily available for us to do…

> And over time, the technology to do this will become cheap and readily available for us to do so. But then the English models will be even better and you'll be back to square one. My guess is that things are going to become more and more American. If you assume that "culture" is a resource like "microchips", then from economic point of view it makes sense to have one country specialize in producing it, and the rest…

> If you assume that "culture" is a resource like "microchips"

I do not. American culture exports American values, which are not universal. Simplest examples being the attitudes towards violence and nudity, which are very different in Europe, and vary within Europe as well.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#202
post #17

Earlier quoted context omitted.

Exactly, if there's one thing transformers are good at it's translation. One I've found particularly nice: any question ChatGPT can answer in English it can answer in French. I'm assuming Norwegian too. So there's no point.

Model can speak Lithuanian too, but with a Russian accent which is a big taboo for us.

I wonder, can you ask it not to do that first, and check if that still happens?

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#203

Earlier quoted context omitted.

If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

I task GPT/Claude with researching stuff that pertains to very specific cultural or legal aspects in French politics, on a daily basis. Even though French is a way more common language globally than Norwegian, these models still haven't figured out that, no matter the language I myself speak to them (German or English depending on my mood) their web searches need to be done in French to return reasonable results. I h…

I have the opposite problem. I often have to ask ChatGPT about things related to Norway and I have to constantly correct it when it keeps switching to responding in Norwegian no matter how many times I tell it to only answer in Norwegian when I request it.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#204
post #101
post #86

> Marius Husnes, the Head of IT Platform at the library (Nasjonlbiblioteket) discussed the project at Huawei’s ID Forum 2026 in Paris, saying that no commercial LLM provider was developing a local (Norwegian) language LLM. He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about…

When I am chatting with ChatGPT - it is fairly obvious that it is American - its native language, its style, its attitude is American - even if we chat in Danish. Just as we cannot rely on Netflix and HBO to produce Scandinavian TV-shows even though they might do at the moment, we need to make our own stuff in this area too. And over time, the technology to do this will become cheap and readily available for us to do…

I chat to it in English instead of my native Spanish not only because of performance, but because I cannot stand the unnatural style it has in Spanish.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#205

Earlier quoted context omitted.

> And over time, the technology to do this will become cheap and readily available for us to do so. But then the English models will be even better and you'll be back to square one. My guess is that things are going to become more and more American. If you assume that "culture" is a resource like "microchips", then from economic point of view it makes sense to have one country specialize in producing it, and the rest…

> If you assume that "culture" is a resource like "microchips" I do not. American culture exports American values, which are not universal. Simplest examples being the attitudes towards violence and nudity, which are very different in Europe, and vary within Europe as well.

Which is already changing thanks to the American influence.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#206
post #63

Earlier quoted context omitted.

I have no idea if the similar spelling will somehow help - I used that mostly because it's a simple way if illustrating the close relationship, but I suspect you'd find that the meanings of closely related words are likely to more directly overlap. The grammar is perhaps more likely to help. Similar word order etc. Even weirdness like German - my only top grade on a German essay in school was one where I on purpose i…

The same thing works for guessing German grammar from English. The farther back you go in English, the more its grammar resembles German. "What sayest thou?" -> "Was sagst du?" In fact, for the above, you don't even have to know a single German word. You just have to know what for question words, "wh" -> "w", that the English "y" at the end of a syllable usually comes from an older Germanic "g" sound, and that "th" w…

That's interesting. I haven't thought about it in that direction before. I'm "of course" aware of the High German consonant shift, which also muddled things a lot (the continuum around to North Sea is a lot "cleaner" if you look at Plattdeutsch instead), but never thought much about what other simple transformations to apply with standard modern German.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#207
post #186

Earlier quoted context omitted.

Oddly enough, my wife was recently involved in a project to translate historical crime novels from Norwegian; since all the available late 20th century Scandinavian crime novels have already been translated and turned into popular TV series, the plan was to go further back. Into the 1930s. The first cut was done with LLMs, but encountered the problem that (a) Norwegian itself has changed noticeably since then, in bot…

Sorry if I was unclear, I didn't want to give the impression I think translations or even transcriptions in some cases is easy, or without problems, or not painstakingly time-consuming, it very much is. I just think building a LLM from scratch is ever harder, with more potential problems that are harder to solve, more time-consuming and even more resource-intensive.

It would require an investment, but those will pay dividends later, as it becomes easier to train LLMs on/for Norwegian. If we need to translate everything to English we might as well just drop using Norwegian altogether. Practically everyone speaks English fluently already...

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#208

Earlier quoted context omitted.

Sorry if I was unclear, I didn't want to give the impression I think translations or even transcriptions in some cases is easy, or without problems, or not painstakingly time-consuming, it very much is. I just think building a LLM from scratch is ever harder, with more potential problems that are harder to solve, more time-consuming and even more resource-intensive.

It would require an investment, but those will pay dividends later, as it becomes easier to train LLMs on/for Norwegian. If we need to translate everything to English we might as well just drop using Norwegian altogether. Practically everyone speaks English fluently already...

> as it becomes easier to train LLMs on/for Norwegian

Why would it be easier in the future? The advances we see with LLMs today require a huge amount of data, and it's getting hard getting the amount of data just using any language, I'm having a hard time seeing how it'd get easier for Norwegians to build their own LLM, unless they seriously start to ramp up how much Norwegian content they're putting out.

> If we need to translate everything to English we might as well just drop using Norwegian altogether. Practically everyone speaks English fluently already...

Yeah I mean with that black and white perspective you can pretty much do anything and it won't matter for anything :) I think for the rest of us, what we speak daily and what we rely on professionally, can differ, and that's OK. But maybe this is just my broken Swedish mind being so used to using English professionally but then conversing in Spanish outside of work daily, YMMV.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#210
post #209

Norway isn't in the EU (no restrictions on Huawei) and has cheap electricity, could become an ai powerhouse.

Norway is in the EEA, which have adopted the existing EU Cybersecurity Act.

Thank you, I didn't know that.
Post reply on HN