Live data from Hacker News

Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

github.com

141–150 of 261 posts

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#141

Earlier quoted context omitted.

That’s the joke.

It isn't. The entirety of the comment I responded to is "Oh no, someone is profiting off of their work without proper attribution!?!?" It's a valid point, but references someone using content created by others for profit. I'm objecting to equating this project with the work done by the original content creators. They're not remotely the same thing. I understand how the internet works and how people respond to others…

It's time to stop digging

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#142
post #92

Earlier quoted context omitted.

I like the [dead] comment theory that they proposed a huge LLM training budget to the government, kept most of the money, and released a cheap merge to justify the grift.

It's kinda weird to claim extraordinary results in such case though, as that brings a lot of eyes to it.

Nothing weird. The mayor wanted something brag about. That Rio, my friend.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#143

Earlier quoted context omitted.

It is a recurrent Brazilian meme: Rio is known in Brazil as "terra de bandido" (gangster's land). The majority of their politicians have ties to organized crime. There is a virtual revolving door between police and crime, where people migrate from one to the other. It is like Chicago in the 20s, Naples and Medelin in the 80s or Moscow and Culiacan (Sinaloa, Mexico) today.

Somehow I doubt that political affiliations with crime syndicates are affecting heavily the dispositions of LLM developers. The industry itself though is one of incest.

Politicians don't come from outer space, they emerge locally and were raised swimming in an imaginary that has normalized the morals that eventually end up expressed at the top.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#144

> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.

> A simple linear combination of every weight did not degrade the performance of the model, but enhanced it. Enhanced it on a couple benchmarks, supposedly. The game is to turn knobs until you get a benchmark run that shows an improvement, then ship it. There are a lot of fine tunes and chimera models on HuggingFace that are supposedly better at some specific test, but when you use them for anything else they're usua…

I don't think your last point is correct. Ablation, when done correctly, seems to increase the quality and typically also the performance too.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#145
post #92

I'm honestly surprised that they even had the inclination to attempt creating a model. I guess it's bullish that a municipal IT department had the guts to try this?

I like the [dead] comment theory that they proposed a huge LLM training budget to the government, kept most of the money, and released a cheap merge to justify the grift.

This would be so very brazilian of them.

Source: am Huelander.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#146

> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.

What I find fascinating is the idea that there might be a set of "secret" tweaks that when applied to those weights (or even smaller models) could result in an intelligence simulation that could vastly surpass even something like Fable.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#147
post #77

Earlier quoted context omitted.

why not?

It is a recurrent Brazilian meme: Rio is known in Brazil as "terra de bandido" (gangster's land). The majority of their politicians have ties to organized crime. There is a virtual revolving door between police and crime, where people migrate from one to the other. It is like Chicago in the 20s, Naples and Medelin in the 80s or Moscow and Culiacan (Sinaloa, Mexico) today.

Rio is kinda funny as a litmus test - federal government creates laws to try and curb some of the corruption, and Rio produces better and better corrupts - so far Rio is winning.

BTW wasn't it a few months ago the current governor wanted to leave to be able to run as a candidate, so he asked a supreme justice to step in in as governor, since there wasn't anyone else that technically could?

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#148
post #46

Earlier quoted context omitted.

I wouldn’t describe what happened here as incompetence. As a “carioca”, I am pleasantly surprised to know that the government’s IT department is involved in AI work — even without the budget to create its own models from scratch.

It is a testament to the bloat and overreach of the Brazilian state in the economy. Such endeavors should be left to the private sector

I disagree. I’d prefer if my government invested more in AI solutions, so as not to depend so much on foreign technology.

In an ideal world, Brazil would have a thriving private sector, capable of competing even in the AI sector. Unfortunately, that’s not the case, and I believe that without government action such endeavors won’t really succeed.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#150
post #12

This is fascinating that it worked though. Can we just merge all the open weight models and get something better?

I imagine it'd work the same as merging all the good-tasting foods to get an even tastier one

[dead]
Post reply on HN