“Well, Steve (Jobs), I think it’s more like we both had this rich neighbor named Xerox, and I broke into his house to steal the TV set, but I found out that you had already stolen it.” -- Bill Gates
Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
61–70 of 261 posts
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#62[deleted]
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#63The model's webpage at https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B says it's a merge now. It previously didn't contain this paragraph: >The model is built via a merge of https://huggingface.co/nex-agi/Nex-N2-Pro and https://huggingface.co/Qwen/Qwen3.5-397B-A17B , proceeded by On-Policy Distillation from a stronger model. We detected an incorrect upload in the previous version, where the base merged versio…
It wasnt framed as an issue which is the norm breakage I think you’re reacting to, as in they didnt ask that the readme be updated etc, but it is common now for folks to use a project’s issue tracker to name and shame them in a place they cant easily ignore.
Whether that’s right, prosocial, or professional is up for debate (as well as if any single definition of etiquette can be expected in 2026 on an issue tracker).
But surely you can see the optics reason why someone would take their complaint to the repo directly? It pressures the maintainers to respond, it allows for a pile on from the internet, and makes any decision to lock down a hostile thread into its own kind of statement.
The maintainers should absolutely post an official response and lock the thread though, it will likely get ugly in there.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#64Oh no, someone is profiting off of their work without proper attribution!?!?
This is an open weights model based on other open weights models. The dispute is that they released it with claims about having done some post training that improved the outputs. It was discovered that the model was not post trained like they claimed. The HF page now says it’s a merge of models, which wasn’t there before. They’re trying to claim they accidentally uploaded the wrong model to HF and that they’ll upload…
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#65“Well, Steve (Jobs), I think it’s more like we both had this rich neighbor named Xerox, and I broke into his house to steal the TV set, but I found out that you had already stolen it.” -- Bill Gates
lmao i really hope this is a real quote cuz it’s a banger
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#66I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#67Earlier quoted context omitted.
Without evidence, your comment is just bad mouthing. I have been involved in academia, including in Brazil, and I don't find academia there any more copycat than any other institution, including top tier ones.
This is very easy to prove [1][2]. Brazil has that reputation in the broarder academic world, and it's for a reason. [1] https://www.sciencedirect.com/science/article/abs/pii/S17511... [2] https://www.scielo.br/j/aac/a/xNytDrrrHdyK4XPcHBRJZmd/?lang=...
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#68Earlier quoted context omitted.
This is an open weights model based on other open weights models. The dispute is that they released it with claims about having done some post training that improved the outputs. It was discovered that the model was not post trained like they claimed. The HF page now says it’s a merge of models, which wasn’t there before. They’re trying to claim they accidentally uploaded the wrong model to HF and that they’ll upload…
How do they just splice two models together?
In the early days of Llama there were a lot of experiments like this. There were even some interesting combinations of models where they stacked layers of different models together or even added more layers with interesting results.
But announcing that you spliced two models together isn't very impressive in 2026, so they announced that they had done their own post training and outdid the big labs. They thought nobody would look close enough to notice.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#69“Well, Steve (Jobs), I think it’s more like we both had this rich neighbor named Xerox, and I broke into his house to steal the TV set, but I found out that you had already stolen it.” -- Bill Gates
> Bill Gates had somehow manifested, alone, surrounded by ten Apple employees. … Steve started yelling at Bill, asking him why he violated their agreement.
And what’s more interesting is the conclusion:
> Apple filed a monumental copyright lawsuit against Microsoft in 1988, but they eventually lost on a technicality (the judge ruled that Apple inadvertently gave Microsoft a perpetual license to the Mac user interface in November 1985).
Microsoft didn’t steal Apple’s GUI … Apple gave it to them.
Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model
#70Earlier quoted context omitted.
Unlike the big companies who do this, which often are merely impure scams on tax payer money a little more downstream.
Great, now we're defending embezzlement and fraud with public funds on HN, because we really really hate big business. A child caught doing something bad will cry "but my friends also did it!", is that the level of reasoning hackers want to be at?