Earlier quoted context omitted.
You'd think so. It seems like there are a lot of odd gaps like that. I also have a favourite English language PhD thesis I ask every new model about that they still struggle to find even though there's a Wikipedia article about it that links a blog post I wrote about it. Anyone who thinks they've exhausted even publicly crawlable resources should ask them about some obscure stuff.
you might be surprised if you take this approach.. give key words and phrases in small amounts, each sentence of a prompt building on a previous sentence. Take a an example that is not very hard, like Lewis Carrol Alice in Wonderland original text. Although a quick question might get things sort of wrong, or miss details, if you guide the LLM to a certain part of the story, then a certain set of characters in that pa…
Norway's 2 petabytes of Huawei flash storage and LLM training
191–200 of 227 posts
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#192Earlier quoted context omitted.
> high quality transcriptions and translations of the stories currently described only in Norwegian into English You make it sound like an easier task than training an LLM. I'd argue it's not obvious, and would assume the contrary.
Yes, why wouldn't it be easier to transcribe and translate, skills humanity had for centuries, compared to LLMs that we've only learnt to build these last few years, and even require a frikken computer to do? Of course one of these is harder than the other...
With absolutely no insight into why, which one has better odds to happen first is obvious to me.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#193Earlier quoted context omitted.
The point is that norway willl have its own LLM. And will not have dependencies to another state or private company. The goal is not to be the best model. But to have a model that include more Norwegian data then other LLM and that it's not screwed against other sources.
But what does that give you? If the model is far less capable? What will it do for you with that Norwegian data, that a better model could not do with better search or context?
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#194Earlier quoted context omitted.
Yes, why wouldn't it be easier to transcribe and translate, skills humanity had for centuries, compared to LLMs that we've only learnt to build these last few years, and even require a frikken computer to do? Of course one of these is harder than the other...
Look at it from this lens: translating and transcribing these stories hasn't happened for the centuries they existed, while as you point out the skills where always there. In contrast LLMs have been here for a few years at most and everyone and their dogs are trying to get in on the "race". With absolutely no insight into why, which one has better odds to happen first is obvious to me.
Having insights into both translations, transcriptions and attempting to build LLMs myself, I'm fairly sure which effort would be successful first, regardless of how many attempt it first.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#195How true is this statement: "He asserted that any country with its own language that did not have a sovereign LLM trained in that language was at a disadvantage as a globally trained, English-speaking LLM would not know about that country’s history, news and culture that was described in the local language." I thought all big players already train on basically everything remotely available to them no matter the langu…
If you want LLMs to have knowledge of the Norwegian language, wouldn't the most obvious thing to do be to build a good training dataset and make the dataset widely available? Why go to the expense of training your own model, especially when it will be inferior to state of the art models.
Uuh.. No? Especially of the training data, as in this case, is of better quality.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#196Earlier quoted context omitted.
He’s right though, although it’s not entirely about the training corpus. It’s about the tokenizer that tokenizes substrings more efficiently based on a necessary bias towards a target language. English oriented LLMs are more powerful for English than other languages because the token space is more parsimonious in English language. Try any online Anthropic tokenizer that calls their api with common English words (typi…
Did you even try to verify your claims. I tested it on few translations on wikipedia articles using [1] and it takes 15-20% more tokens for Norwegian. English performs the best because there is more data in English and high quality sources are either only in English or there is a good translation in English. [1]: https://platform.openai.com/tokenizer
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#197> The Olivia system is an HPE Cray Supercomputing EX system, with 448 GPUs and 64,512 CPU cores. Training a sovereign LLM with this meager hardware as opposed to a LORA on some open source model seems like a huge mistake and a potential red flag. There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Which…
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#198Earlier quoted context omitted.
He’s right though, although it’s not entirely about the training corpus. It’s about the tokenizer that tokenizes substrings more efficiently based on a necessary bias towards a target language. English oriented LLMs are more powerful for English than other languages because the token space is more parsimonious in English language. Try any online Anthropic tokenizer that calls their api with common English words (typi…
Did you even try to verify your claims. I tested it on few translations on wikipedia articles using [1] and it takes 15-20% more tokens for Norwegian. English performs the best because there is more data in English and high quality sources are either only in English or there is a good translation in English. [1]: https://platform.openai.com/tokenizer
https://www.google.com/search?q=tokenizer+efficiency+by+languageRe: Norway's 2 petabytes of Huawei flash storage and LLM training
#199may not be the most efficient way to go about things, but there remains a seemingly obvious use case for non-latin languages to do things from scratch. see sarvam.ai and their tokenisation improvements on local languages [1]. not every llm needs to help with coding, nor it needs to already become Babel fish. language is culture, so i can see the motivation behind their initiative. it must be nice to afford to do this…
>but there remains a seemingly obvious use case for non-latin languages to do things from scratch >see sarvam.ai and their tokenisation improvements on local languages You don't need to build from scratch to improve tokenization, though. Russia's T-Bank was able to increase generation speeds by 1.5-3x by changing a stock Qwen's tokenizer to include 5 times more Cyrillic tokens (+ post-training on a Russian corpus).
the great thing about the current momentum is that someone can test this hypothesis by applying the T-Bank approach to the same set of languages and compare outcomes.
unfortunately not everyone has the same level of respectable compute this easily available. at least those outside of the ZIRP/VC ecosystem of the valley.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#200Even entire governments are captured by a mild LLM psychosis. Which is sad in the case of Norway. I lived in Norway for two years and always found their government to be highly rational, this is not a rational use of public funds (but I suppose they have plenty of capital). Western society is completely captured by this form of psychosis and its going to bite us in the a* very soon. I firmly believe all the Boomer le…
I think it is highly rational. You see it from the wrong point of view. It seems to be less a short utilitarian project or economic endeavour, but a cultural one. Think about it more of in terms of applied humanities. Which languages go extinct, which cultures disappear and are superseded by a monocultural globalist hegemony.