Earlier quoted context omitted.
Even if the model will reply in my language, I often notice it searching in english. Or thinking in english. There's always something lost in translation. Sometimes it's just minor nuances. Other times it mangles the legal facts with those of other countries.
This sounds like the problem of people calling "911" as the emergency number which they see in so much US-American media but which is not the emergency number in their own country.
Norway's 2 petabytes of Huawei flash storage and LLM training
211–220 of 227 posts
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#212Seems like they should be building an MCP service rather than training an entirely new LLM...
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#213Earlier quoted context omitted.
They made the cultural case, you have no idea how strong this is in places like quebec, nordics, france, russia etc
Can confirm that. Norway may have a small population, but if you live there you'll think it's truly the center of the world (aside from the US. Norwegians love America)
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#214Earlier quoted context omitted.
> There is no way these people have the resources to train a fully fledged LLM, so claiming that is their goal makes me think they don't intend for the LLM to be useful. Depends on what they are doing and why. but at most big labs, only the final model training happens on the big clusters. a lot of experimentation happens on So for fast iteration, this seems fine.
This is the use case for the small NVIDIA boxes that a researcher can have on their desk for $5k and do useful experiments before spending all the grant money on a huge training run for the final product.
but that only gets you so far, you need bigger multi-GPU setup to do the higher dimension stuff. You can use a DGX, but again thats limiting up to a certain point.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#215Earlier quoted context omitted.
It's really fantastic. I just wished there were fewer restrictions on the content that is accessible. (a lot is only accessible from Norwegian IP addresses, so it's one of the main reasons I maintain a VPN as I'm Norwegian but live in the UK; a second set is only available from the IP addresses of libraries or research institutions - still huge amounts that are generally available, though)
My biggest gripe with it are the restrictions, indeed. When searching through the closed newspapers, you have to apply for access manually, which gives you 8 hours of access. Great. Only that the access is seemingly manually granted - so if you apply 16:05 on a Friday, chances are you won't get any access until 9-10 the next Monday. With that said, I do understand why it is like that. If people could apply via API, a…
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#216Earlier quoted context omitted.
It would require an investment, but those will pay dividends later, as it becomes easier to train LLMs on/for Norwegian. If we need to translate everything to English we might as well just drop using Norwegian altogether. Practically everyone speaks English fluently already...
> as it becomes easier to train LLMs on/for Norwegian Why would it be easier in the future? The advances we see with LLMs today require a huge amount of data, and it's getting hard getting the amount of data just using any language, I'm having a hard time seeing how it'd get easier for Norwegians to build their own LLM, unless they seriously start to ramp up how much Norwegian content they're putting out. > If we nee…
Furthermore one can reuse investments in data (both agreements, infrastructure and datasets), compute (GPUs, servers) and know-how (training scripts, experienced engineers).
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#217Earlier quoted context omitted.
> as it becomes easier to train LLMs on/for Norwegian Why would it be easier in the future? The advances we see with LLMs today require a huge amount of data, and it's getting hard getting the amount of data just using any language, I'm having a hard time seeing how it'd get easier for Norwegians to build their own LLM, unless they seriously start to ramp up how much Norwegian content they're putting out. > If we nee…
These models will never compete with frontier models and do not need to - it is about hitting a good-enough, not being the best. Behind the frontier, getting to a certain performance level, is getting easier over time - both sample and compute efficiency is going up. Furthermore one can reuse investments in data (both agreements, infrastructure and datasets), compute (GPUs, servers) and know-how (training scripts, ex…
I understand and agree building the LLMs yourself comes with more benefits, long-term ones especially, but still it's harder, more expensive and really time consuming work.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#218Earlier quoted context omitted.
These models will never compete with frontier models and do not need to - it is about hitting a good-enough, not being the best. Behind the frontier, getting to a certain performance level, is getting easier over time - both sample and compute efficiency is going up. Furthermore one can reuse investments in data (both agreements, infrastructure and datasets), compute (GPUs, servers) and know-how (training scripts, ex…
But are you seriously under the belief that all of that, plus all the other things you're forgetting about, is easier, cheaper and faster than transcriptions and translations? I understand and agree building the LLMs yourself comes with more benefits, long-term ones especially, but still it's harder, more expensive and really time consuming work.
But for a national lab I think it is money well spent to figure out the possibilities and limitations of a native-language LLMs for languages with order of 5M-10M speakers.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#219Earlier quoted context omitted.
It's also a bit funny because Norway definitely has enough money to hire a team of Anthropic's best to go out there and train them a model that does whatever they want. They probably have enough money to fund their own Anthropic competitor.
>They probably have enough money to fund their own Anthropic competitor. Which is bizarre to me Norway doesn't have a booming tech sector with all hat wealth fund acting as the biggest VC. They instead use their wealth fund to invest in US's tech sector. Baffling.
Re: Norway's 2 petabytes of Huawei flash storage and LLM training
#220Earlier quoted context omitted.
Yeah, was about to comment that too, instead of training a new model and new weights exclusively for Norwegian (and expecting/wanting every other small/medium-sized country to do the same) which seems infinity harder, they could have made high quality transcriptions and translations of the stories currently described only in Norwegian into English, and making it all public. I guess there still would be a worry that i…
Oddly enough, my wife was recently involved in a project to translate historical crime novels from Norwegian; since all the available late 20th century Scandinavian crime novels have already been translated and turned into popular TV series, the plan was to go further back. Into the 1930s. The first cut was done with LLMs, but encountered the problem that (a) Norwegian itself has changed noticeably since then, in bot…
Nynorsk and bokmål is not dialects but variants of written Norwegian.