Live data from Hacker News

Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

github.com

101–110 of 261 posts

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#103

Earlier quoted context omitted.

How do they just splice two models together?

Out of curiosity, how was it discovered? You would have to look for it to find this linear combination.

Without the system prompt, asking its name results in it responding with the name of the model they're ripping from. That would certainly draw your eyes to the right places.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#104
post #101

Earlier quoted context omitted.

This is a pure scam on tax payer money. But what else would be expected?

Apparently no public money was involved.

This is contrary to the mayor's words on Twitter.

> An open AI model trained in Rio with public funding over the last year by @Prefeitura_Rio surpassing all other models.

https://x.com/CavaliereRio/status/2065984620626129026

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#105

> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.

> A simple linear combination of every weight did not degrade the performance of the model, but enhanced it. Enhanced it on a couple benchmarks, supposedly. The game is to turn knobs until you get a benchmark run that shows an improvement, then ship it. There are a lot of fine tunes and chimera models on HuggingFace that are supposedly better at some specific test, but when you use them for anything else they're usua…

They seem to have deleted most of the README now, but the archived version has benchmarks.

https://web.archive.org/web/20260614082641/https://huggingfa...

And the Nex benchmarks for comparison

https://huggingface.co/nex-agi/Nex-N2-Pro

Rio seems to be about halfway between Qwen 3.5 and Nex, as you'd expect?

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#106

Earlier quoted context omitted.

Attribution isn't the relevant part. Lying about your lab's capabilities is.

That's also something all the AI companies have been doing.

They’re using public money to “train” this.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#107
post #103

Earlier quoted context omitted.

Out of curiosity, how was it discovered? You would have to look for it to find this linear combination.

Without the system prompt, asking its name results in it responding with the name of the model they're ripping from. That would certainly draw your eyes to the right places.

Why is this? Do labs reinforce the model name during training? I was under the impression that this sort of "self-knowledge" always came from the system prompt, but I guess not...

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#108
It's absolutely insane to me that we are now at a point where the top of the front page of hacker news is a random GitHub issue about attribution to some random LLM merge, written in just the most disgusting AI slop style.

I would like to downvote this please.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#109

> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.

https://thickets.mit.edu

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#110

> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.

It shows that LLMs are an extremely wasteful approach to intelligence.

[deleted]
Post reply on HN