Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
1–10 of 55 posts
Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#2And how did they decide that, e.g., German or Dutch would make the model worse?
Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#3Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#4Is there something similar for romance or Germanic languages? And how did they decide that, e.g., German or Dutch would make the model worse?
Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#5I have had this question. How much better would common LLMs (Llama, GPTN) be if they were only trained in one language? I have to assume they would perform better, but I might be wrong.
Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#6I have had this question. How much better would common LLMs (Llama, GPTN) be if they were only trained in one language? I have to assume they would perform better, but I might be wrong.
Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#7I hope to see this used to generate a customized curriculum for each neurodiverse child so that we can live in a more equitable society.
Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#8Is there something similar for romance or Germanic languages? And how did they decide that, e.g., German or Dutch would make the model worse?
I don't think they decided that, they included Finish which is completely unrelated to the other nordic languages. If they just picked languages that are related for cross learning including Dutch or German would have made more sense indeed.
Re: Viking 7B: open LLM for the Nordic languages trained on AMD GPUs
#9I have had this question. How much better would common LLMs (Llama, GPTN) be if they were only trained in one language? I have to assume they would perform better, but I might be wrong.
Just like adding code to textual models helps the model develop its reasoning capabilities, it seems like adding more languages helps in other areas too. What is needed is more good quality data to train on...
These architectures are less capable than brains in many ways. So, we should expect them to have such trade-offs. An efficient one should work fine on English, mathematical notation, and a programming language. Maybe samples of others that illustrate unique concepts. I’m also curious how many languages or concepts you can add to a given architecture before its effectiveness starts dropping.