Oh no, someone is profiting off of their work without proper attribution!?!?
This is a pure scam on tax payer money. But what else would be expected?
Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
101–110 of 261 posts
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#102Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#103Earlier quoted context omitted.
How do they just splice two models together?
Out of curiosity, how was it discovered? You would have to look for it to find this linear combination.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#104Earlier quoted context omitted.
This is a pure scam on tax payer money. But what else would be expected?
Apparently no public money was involved.
> An open AI model trained in Rio with public funding over the last year by @Prefeitura_Rio surpassing all other models.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#105> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.
> A simple linear combination of every weight did not degrade the performance of the model, but enhanced it. Enhanced it on a couple benchmarks, supposedly. The game is to turn knobs until you get a benchmark run that shows an improvement, then ship it. There are a lot of fine tunes and chimera models on HuggingFace that are supposedly better at some specific test, but when you use them for anything else they're usua…
https://web.archive.org/web/20260614082641/https://huggingfa...
And the Nex benchmarks for comparison
https://huggingface.co/nex-agi/Nex-N2-Pro
Rio seems to be about halfway between Qwen 3.5 and Nex, as you'd expect?
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#106Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#107Earlier quoted context omitted.
Out of curiosity, how was it discovered? You would have to look for it to find this linear combination.
Without the system prompt, asking its name results in it responding with the name of the model they're ripping from. That would certainly draw your eyes to the right places.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#108I would like to downvote this please.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#109> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#110> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.
It shows that LLMs are an extremely wasteful approach to intelligence.