Live data from Hacker News

Norway's 2 petabytes of Huawei flash storage and LLM training

blocksandfiles.com

191–200 of 227 posts

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#191
post #66

Earlier quoted context omitted.

You'd think so. It seems like there are a lot of odd gaps like that. I also have a favourite English language PhD thesis I ask every new model about that they still struggle to find even though there's a Wikipedia article about it that links a blog post I wrote about it. Anyone who thinks they've exhausted even publicly crawlable resources should ask them about some obscure stuff.

you might be surprised if you take this approach.. give key words and phrases in small amounts, each sentence of a prompt building on a previous sentence. Take a an example that is not very hard, like Lewis Carrol Alice in Wonderland original text. Although a quick question might get things sort of wrong, or miss details, if you guide the LLM to a certain part of the story, then a certain set of characters in that pa…

For the PhD thesis in question, I've actually tested a lot of requests about different parts of it, and both Claude and ChatGPT still draws a total blank if you don't let them do searches.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#192

Earlier quoted context omitted.

> high quality transcriptions and translations of the stories currently described only in Norwegian into English You make it sound like an easier task than training an LLM. I'd argue it's not obvious, and would assume the contrary.

Yes, why wouldn't it be easier to transcribe and translate, skills humanity had for centuries, compared to LLMs that we've only learnt to build these last few years, and even require a frikken computer to do? Of course one of these is harder than the other...

Look at it from this lens: translating and transcribing these stories hasn't happened for the centuries they existed, while as you point out the skills where always there. In contrast LLMs have been here for a few years at most and everyone and their dogs are trying to get in on the "race".

With absolutely no insight into why, which one has better odds to happen first is obvious to me.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#193
post #140
post #26

Earlier quoted context omitted.

The point is that norway willl have its own LLM. And will not have dependencies to another state or private company. The goal is not to be the best model. But to have a model that include more Norwegian data then other LLM and that it's not screwed against other sources.

But what does that give you? If the model is far less capable? What will it do for you with that Norwegian data, that a better model could not do with better search or context?

A model has many dimensions. You can't have them on one scale from good to bad. The model will most likely be poor at coding. But will give better answers about Norwegian cultur. I assume the tone of voice will be (by default) much closer to how Norwegians talk and write then what we current see from model from the US. They seem to be a bit to much.. Norwegian people are a bit more down to earth

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#194

Earlier quoted context omitted.

Yes, why wouldn't it be easier to transcribe and translate, skills humanity had for centuries, compared to LLMs that we've only learnt to build these last few years, and even require a frikken computer to do? Of course one of these is harder than the other...

Look at it from this lens: translating and transcribing these stories hasn't happened for the centuries they existed, while as you point out the skills where always there. In contrast LLMs have been here for a few years at most and everyone and their dogs are trying to get in on the "race". With absolutely no insight into why, which one has better odds to happen first is obvious to me.

Sure, it isn't as "hot" to translate stuff as it used to be some hundreds of years ago, and building LLMs surely is "hot" today, I don't doubt more people are attempting to build LLMs today than translating huge datasets, especially if we narrow the two to exclusively "In Norwegian".

Having insights into both translations, transcriptions and attempting to build LLMs myself, I'm fairly sure which effort would be successful first, regardless of how many attempt it first.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#195

How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…

If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

> Why go to the expense of training your own model, especially when it will be inferior to state of the art models.

Uuh.. No? Especially of the training data, as in this case, is of better quality.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#196

Earlier quoted context omitted.

He’s right though, although it’s not entirely about the training corpus. It’s about the tokenizer that tokenizes substrings more efficiently based on a necessary bias towards a target language. English oriented LLMs are more powerful for English than other languages because the token space is more parsimonious in English language. Try any online Anthropic tokenizer that calls their api with common English words (typi…

Did you even try to verify your claims. I tested it on few translations on wikipedia articles using [1] and it takes 15-20% more tokens for Norwegian. English performs the best because there is more data in English and high quality sources are either only in English or there is a good translation in English. [1]: https://platform.openai.com/tokenizer

Tests I've done with NO and FI texts, for the same number of characters, with the GPT5 tokenizer I get around 2x the tokens than EN. With the older tokenizers it's more like 2x or even 3x.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#197

> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…

LoRA won't fix the tokenization problem. Norwegian on a typical English-heavy BPE vocab uses 1.5-2x more tokens per word — that compounds into real inference cost, not just quality

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#198

Earlier quoted context omitted.

He’s right though, although it’s not entirely about the training corpus. It’s about the tokenizer that tokenizes substrings more efficiently based on a necessary bias towards a target language. English oriented LLMs are more powerful for English than other languages because the token space is more parsimonious in English language. Try any online Anthropic tokenizer that calls their api with common English words (typi…

Did you even try to verify your claims. I tested it on few translations on wikipedia articles using [1] and it takes 15-20% more tokens for Norwegian. English performs the best because there is more data in English and high quality sources are either only in English or there is a good translation in English. [1]: https://platform.openai.com/tokenizer

Tokenizer efficiency varying by languages, by as much as up to 15x, is very well known and established

  https://www.google.com/search?q=tokenizer+efficiency+by+language

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#199
post #185

may not be the most efficient way to go about things, but there remains a seemingly obvious use case for non-latin languages to do things from scratch. see sarvam.ai and their tokenisation improvements on local languages [1]. not every llm needs to help with coding, nor it needs to already become Babel fish. language is culture, so i can see the motivation behind their initiative. it must be nice to afford to do this…

>but there remains a seemingly obvious use case for non-latin languages to do things from scratch >see sarvam.ai and their tokenisation improvements on local languages You don't need to build from scratch to improve tokenization, though. Russia's T-Bank was able to increase generation speeds by 1.5-3x by changing a stock Qwen's tokenizer to include 5 times more Cyrillic tokens (+ post-training on a Russian corpus).

the improvements for sarvam was with the amount of tokens used to represent words in english vs non-english languages.

the great thing about the current momentum is that someone can test this hypothesis by applying the T-Bank approach to the same set of languages and compare outcomes.

unfortunately not everyone has the same level of respectable compute this easily available. at least those outside of the ZIRP/VC ecosystem of the valley.

Re: Norway's 2 petabytes of Huawei flash storage and LLM training

#200
post #64

Even entire governments are captured by a mild LLM psychosis. Which is sad in the case of Norway. I lived in Norway for two years and always found their government to be highly rational, this is not a rational use of public funds (but I suppose they have plenty of capital). Western society is completely captured by this form of psychosis and its going to bite us in the a* very soon. I firmly believe all the Boomer le…

I think it is highly rational. You see it from the wrong point of view. It seems to be less a short utilitarian project or economic endeavour, but a cultural one. Think about it more of in terms of applied humanities. Which languages go extinct, which cultures disappear and are superseded by a monocultural globalist hegemony.

Exactly. Nasjonalbiblioteket (National Library of Norway) has centuries of written material (Bokmål, Nynorsk and some Sami) and decades of audio and video material featuring varied dialects from all over the country. I believe training models that encompass this information can help in preserving both our language, history, and culture for future generations that increasingly turn to AI to get their information.
Post reply on HN