Live data from Hacker News

Apertus – Open Foundation Model for Sovereign AI

apertvs.ai

31–40 of 198 posts

Re: Apertus – Open Foundation Model for Sovereign AI

#31
post #10

For a model that claims to focus on many languages, it's quite unreliable when it comes to simple questions like "how to say X in language Y" or "how to conjugate verb X in language Y". It keeps hallucinating words that do not exist, and when corrected, it only hallucinates a new lie.

it probably doesnt know what language each set of words is referencing.

i doubt they are including a lot of training data labeled with the language.

"how to say X in language Y" is a different task from saying X in language Y

Re: Apertus – Open Foundation Model for Sovereign AI

#33
post #20

It's good that there is a movement for open LLMs, but it's not where the battleground is right now . The battleground is local vs service LLMs, and we are losing that battle badly despite all the software being here now and viable, entirely because UX sucks. How many normal people do you know who use "ChatGPT"? A lot, probably. How many even know what "Gemma" is, let alone have downloaded llama.cpp, a GGUF file from…

> We are sleepwalking into slavery.

That’s a bit hyperbolic…

Re: Apertus – Open Foundation Model for Sovereign AI

#34
post #20

It's good that there is a movement for open LLMs, but it's not where the battleground is right now . The battleground is local vs service LLMs, and we are losing that battle badly despite all the software being here now and viable, entirely because UX sucks. How many normal people do you know who use "ChatGPT"? A lot, probably. How many even know what "Gemma" is, let alone have downloaded llama.cpp, a GGUF file from…

"Normal people" have never bothered to host their own: photos, music, videos, documents, comunications, etc. To the point that for many their computer is essentially a thin client into someone else's server. Why would we think this same people would care about "personal" inference?

Re: Apertus – Open Foundation Model for Sovereign AI

#35
post #20

It's good that there is a movement for open LLMs, but it's not where the battleground is right now . The battleground is local vs service LLMs, and we are losing that battle badly despite all the software being here now and viable, entirely because UX sucks. How many normal people do you know who use "ChatGPT"? A lot, probably. How many even know what "Gemma" is, let alone have downloaded llama.cpp, a GGUF file from…

it's funny because i made this thing (called enough) that aims to make it easy for non-technical people to get up and running with local models quickly, but it is impossible to figure out how to break through the noise. every thread and comment like this breaks my heart a lil bit

Re: Apertus – Open Foundation Model for Sovereign AI

#36
post #20

It's good that there is a movement for open LLMs, but it's not where the battleground is right now . The battleground is local vs service LLMs, and we are losing that battle badly despite all the software being here now and viable, entirely because UX sucks. How many normal people do you know who use "ChatGPT"? A lot, probably. How many even know what "Gemma" is, let alone have downloaded llama.cpp, a GGUF file from…

Why do you feel the important part _now_ is where the weights get run?

I can see this as a future battleground but access to frontier models (which you cannot run locally) seems a lot more relevant today.

Re: Apertus – Open Foundation Model for Sovereign AI

#37
post #13

Great to see more fully open LLMs. I think a problem with open-weight models is that while you can improve them, you are not going to create the next generation of LLMs by fine-tuning. We are at the mercy of frontier labs for access to SOTA LLMs. For example, Anthropic recently started requiring identity verification for Claude [0], same for OpenAI [1]. If one day China's distillation labs stop releasing their LLMs a…

> China's distillation labs This notion that Chinese labs are merely distilling frontier models is quite an unwarranted slur. Those labs have published WAY more useful research than US labs on RL techniques, novel model architectures, training pipelines, etc. They have also hit intelligence-per-parameter densities that US labs have yet to attain. Apart from that, merely training a model on outputs from another model,…

But have they? I understand that the Chinese side is illuminated and the American side is dark. I disagree that the Chinese labs have created anything that isn't in an American research lab or production dc. Sure the Chinese have published their findings and not for nothing. But are they novel? Unlikely imo

Re: Apertus – Open Foundation Model for Sovereign AI

#38
post #13

Great to see more fully open LLMs. I think a problem with open-weight models is that while you can improve them, you are not going to create the next generation of LLMs by fine-tuning. We are at the mercy of frontier labs for access to SOTA LLMs. For example, Anthropic recently started requiring identity verification for Claude [0], same for OpenAI [1]. If one day China's distillation labs stop releasing their LLMs a…

> China's distillation labs This notion that Chinese labs are merely distilling frontier models is quite an unwarranted slur. Those labs have published WAY more useful research than US labs on RL techniques, novel model architectures, training pipelines, etc. They have also hit intelligence-per-parameter densities that US labs have yet to attain. Apart from that, merely training a model on outputs from another model,…

I recently watched a video for one of these “Chinese Models” it kept insisting it was Claude when the user asked. Sorry, there’s no “slur” here but legit suspicion.

Re: Apertus – Open Foundation Model for Sovereign AI

#40
post #38

Earlier quoted context omitted.

> China's distillation labs This notion that Chinese labs are merely distilling frontier models is quite an unwarranted slur. Those labs have published WAY more useful research than US labs on RL techniques, novel model architectures, training pipelines, etc. They have also hit intelligence-per-parameter densities that US labs have yet to attain. Apart from that, merely training a model on outputs from another model,…

I recently watched a video for one of these “Chinese Models” it kept insisting it was Claude when the user asked. Sorry, there’s no “slur” here but legit suspicion.

https://blog.kilo.ai/p/did-claude-opus-48-distill-alibabas

it happens to all models…when the internet is increasingly generated, things happen

Post reply on HN