Live data from Hacker News

GPT‑NL: a sovereign language model for the Netherlands

tno.nl

141–150 of 325 posts

Re: GPT‑NL: a sovereign language model for the Netherlands

#141
post #70

I keep seeing these "sovereign" LMs time and time again. In Sweden we had GPT-SW3 ( https://www.ai.se/en/project/gpt-sw3 ) and same story there. Instead of burning money on "sovereign" claims, national research labs should instead focus on building on top of solid baselines (like Qwen/Kimi) and finetuning frontier models with real agentic utility that can be applied across actual use cases and can be widely used by i…

Kimi and Qwen come out of China, which means that their training material may be biased e.g. relating to Taiwan [1]. In addition, there is no way to determine what input went into the training, if it was properly licensed, if it was legal (e.g. not contaminated by CSAM), or how the human component of RLHF was sourced - in US models, for example, stories about exploitation like [2] have been floating for years. Assumi…

> Most LLMs focus on the English, German, French and Chinese languages, but everything else is... left behind at best.

that is not true, so please read before make an opinion. The French Mistral project shipped seven+ years ago with 140 languages for example.. language translation was the first LLM task from 2015

Re: GPT‑NL: a sovereign language model for the Netherlands

#142
post #70

I keep seeing these "sovereign" LMs time and time again. In Sweden we had GPT-SW3 ( https://www.ai.se/en/project/gpt-sw3 ) and same story there. Instead of burning money on "sovereign" claims, national research labs should instead focus on building on top of solid baselines (like Qwen/Kimi) and finetuning frontier models with real agentic utility that can be applied across actual use cases and can be widely used by i…

Do we know for sure how much national corpus of knowledge (like dutch) goes into these "global" models and how that affects "localized" model biases? What's wrong with specialized models?

Re: GPT‑NL: a sovereign language model for the Netherlands

#143
post #117

Earlier quoted context omitted.

Kimi and Qwen come out of China, which means that their training material may be biased e.g. relating to Taiwan [1]. In addition, there is no way to determine what input went into the training, if it was properly licensed, if it was legal (e.g. not contaminated by CSAM), or how the human component of RLHF was sourced - in US models, for example, stories about exploitation like [2] have been floating for years. Assumi…

Uh, some would say it's easy to determine what input went into the training for kimi and qwen.. since they were caught stealing it from American labs. Some cultural cliches may never change.

It's well-known that all commercial models are based on stolen content. That doesn't mean there is no filtering/censoring, just that the censoring likely depends on where it's happening…

Re: GPT‑NL: a sovereign language model for the Netherlands

#144

It is crazy that anything Europe gets so much hate. IMO it is important to build models within the boundaries of smaller nations, using their own language. Research has to continue even if it is outside of US and China.

It’s not that it gets hate so much as it’s akin to watching them make announcements that they’re going to make a European google/facebook/tiktok. Sure… they can, except at the end of the day it’s a bit late, regulatory burden will make it comparatively useless, and because of that nobody will ever use it. It will be spending a bunch of taxpayer dollars for press releases. The running joke is that when these “sovereig…

That’s on Wikipedia, it’s not PII, it’s also not going to be relevant to any meaningful IRL work.

I challenge the assumption you can do meaningful work in this field without blatant disregard for intellectual property.

The idea that it’s all down to training size is clearly incorrect, as every expert human learned their craft without nearly the sum total information of the internet. Clearly there are architectural wins to be found.

Besides that, why would everyone just be fine with Opus level AI at best, as that’s all the US is willing to export, and I doubt China will share beyond that.

Sovereign AI is more important than ever after Friday.

Re: GPT‑NL: a sovereign language model for the Netherlands

#145
post #85

Earlier quoted context omitted.

An LLM is an encoding of a culture, a way of viewing the world. They are not neutral technology, they are a direct representation of the training set that has been chosen and how they are aligned. In many ways, they are ideology made code. If we leave building them to the US and China, only their way of seeing things will be digitized. I don't like the idea of that.

Yes and also, US and Chinese models are censored in different ways. US models are way too prudish for personal use in Europe because they're afraid to piss off religious investors. Chinese models are too censored on history and current affairs, eg the tiananmen massacre never happened stuff like that.

Chinese models aren't censored as much as you think, you can download the model and run it somewhere else and they will happily tell you about Tiananmen Square. Or heck, ask DeepSeek via Openrouter, it will do the same.

The censorship works kind of like with Fabel, it kicks in before the model responds.

Re: GPT‑NL: a sovereign language model for the Netherlands

#146

Earlier quoted context omitted.

Kimi and Qwen come out of China, which means that their training material may be biased e.g. relating to Taiwan [1]. In addition, there is no way to determine what input went into the training, if it was properly licensed, if it was legal (e.g. not contaminated by CSAM), or how the human component of RLHF was sourced - in US models, for example, stories about exploitation like [2] have been floating for years. Assumi…

> Most LLMs focus on the English, German, French and Chinese languages, but everything else is... left behind at best. that is not true, so please read before make an opinion. The French Mistral project shipped seven+ years ago with 140 languages for example.. language translation was the first LLM task from 2015

One example is not the same as "most LLMs". My experience is the same with most LLMs. Especially the smaller ones are English oriented (probably makes sense given the size constraints).

Re: GPT‑NL: a sovereign language model for the Netherlands

#147
post #143
post #117

Earlier quoted context omitted.

Uh, some would say it's easy to determine what input went into the training for kimi and qwen.. since they were caught stealing it from American labs. Some cultural cliches may never change.

It's well-known that all commercial models are based on stolen content. That doesn't mean there is no filtering/censoring, just that the censoring likely depends on where it's happening…

> It's well-known that all commercial models are based on stolen content.

Does that mean that Chinese models are the "Robin Hood"s of the AI era?

Re: GPT‑NL: a sovereign language model for the Netherlands

#148

So good to see these developments. Every country should do this. I'd even say every person should gave their own personalized AI running on their own computers. If only the costs involved were not so astronomical.

Why? That doesn't make any sense.

The government would be far better off figuring out how to take commodity models and applying them to government functions where they can, with deterministic scaffolding and guardrails, to make government more efficient, optionally using RL on traces from their use to improve their performance.

Imagine taking models and fine-tuning them / doing RL rollouts to help automate permit application approvals, as applied specifically to Dutch permit processes. That would be a real help to Dutch businesses!

That type of applied AI is more interesting and effective now than just trying to make another foundational model that isn't going to work well or do anything of economic value.

Re: GPT‑NL: a sovereign language model for the Netherlands

#149
post #70

I keep seeing these "sovereign" LMs time and time again. In Sweden we had GPT-SW3 ( https://www.ai.se/en/project/gpt-sw3 ) and same story there. Instead of burning money on "sovereign" claims, national research labs should instead focus on building on top of solid baselines (like Qwen/Kimi) and finetuning frontier models with real agentic utility that can be applied across actual use cases and can be widely used by i…

Kimi and Qwen come out of China, which means that their training material may be biased e.g. relating to Taiwan [1]. In addition, there is no way to determine what input went into the training, if it was properly licensed, if it was legal (e.g. not contaminated by CSAM), or how the human component of RLHF was sourced - in US models, for example, stories about exploitation like [2] have been floating for years. Assumi…

It really doesn't matter if the model sucks and doesn't perform well. Given the funding amount and their lofty ambitions, it seems very unlikely they will be able to pull it off properly.

Yeah China and US models have baises but so will any model. The biases do not get in the way of the product though. You don't open those models just to ask for what happened in Taianaman square or if Taiwan is a state. You dont ask ChatGPT to generate CASM. But they are very good at the tasks you actually expect from a LLM. If you fail at that, nobody will use your model no matter how "ethically sourced" a colonizer-based entity like Europe made it.

Re: GPT‑NL: a sovereign language model for the Netherlands

#150

It is crazy that anything Europe gets so much hate. IMO it is important to build models within the boundaries of smaller nations, using their own language. Research has to continue even if it is outside of US and China.

If a teenager on your street said he was going to spend $1,000 to customize his Honda Civic for his needs, you'd believe him. If he says he's going to build a brand new car, better than a Honda civic, for $10,000, you'd laugh and say good luck.
Post reply on HN