Live data from Hacker News

Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

github.com

51–60 of 261 posts

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#51

Earlier quoted context omitted.

Attribution isn't the relevant part. Lying about your lab's capabilities is.

I do not see anyone lying. The model card says: > Post-trained from Qwen 3.5 397B The model card also says that they use an inference framework based on "SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs" by Shi et al.: https://arxiv.org/abs/2510.05069 So the sources seem properly attributed. They only claim that what they did to "Qwen 3.5 397B" has improved the LLM, including, a…

That's attribution to Qwen team.

There (is/was) no attribution to Nex team (they've released a model based on Qwen 3.5 397B as well).

As per OP link Nex claims that what Rio team released (so far) is just linear interpolation of weights between Nex and OG Qwen model. With no attribution to Nex and zero signs of Rio doing any training of their own.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#52
post #45

Earlier quoted context omitted.

This is a pure scam on tax payer money. But what else would be expected?

Unlike the big companies who do this, which often are merely impure scams on tax payer money a little more downstream.

Great, now we're defending embezzlement and fraud with public funds on HN, because we really really hate big business.

A child caught doing something bad will cry "but my friends also did it!", is that the level of reasoning hackers want to be at?

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#53

Oh no, someone is profiting off of their work without proper attribution!?!?

This is an open weights model based on other open weights models.

The dispute is that they released it with claims about having done some post training that improved the outputs. It was discovered that the model was not post trained like they claimed.

The HF page now says it’s a merge of models, which wasn’t there before. They’re trying to claim they accidentally uploaded the wrong model to HF and that they’ll upload the real one soon.

Basically, they thought they could splice two open weights models together and claim their team had accomplished some amazing post training, but they weren’t smart enough to realize that other researchers would discover that there wasn’t any post training.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#54

[flagged]

Without evidence, your comment is just bad mouthing. I have been involved in academia, including in Brazil, and I don't find academia there any more copycat than any other institution, including top tier ones.

This is very easy to prove [1][2]. Brazil has that reputation in the broarder academic world, and it's for a reason.

[1] https://www.sciencedirect.com/science/article/abs/pii/S17511...

[2] https://www.scielo.br/j/aac/a/xNytDrrrHdyK4XPcHBRJZmd/?lang=...

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#55
post #2

The municipality of Rio de Janeiro (via its IT company IplanRIO) released Rio-3.5-Open-397B, presented as a homegrown Qwen3.5 fine-tune that beats comparable open models on benchmarks. The linked issue argues it's actually a weighted merge of ~60% Nex-N2 Pro + ~40% Qwen3.5-397B-A17B - Nex-N2 having been released about a week earlier.

So the problem isn’t in the missing attribution to Qwen, but with the fact that they didn’t mention Nex-N2 Pro right?

The problem is that they claimed to have made a big achievement with their home grown post training, and they expected to receive a lot of praise for it.

Then researchers looked at the weights and there is no post training at all.

They are now attributing both models they merged, but their excuse for the lack of post training is to claim they accidentally uploaded the wrong files.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#56
post #2

The municipality of Rio de Janeiro (via its IT company IplanRIO) released Rio-3.5-Open-397B, presented as a homegrown Qwen3.5 fine-tune that beats comparable open models on benchmarks. The linked issue argues it's actually a weighted merge of ~60% Nex-N2 Pro + ~40% Qwen3.5-397B-A17B - Nex-N2 having been released about a week earlier.

I didn't know model merging like that was possible. (Obviously possible from a pure software standpoint but I'm surprised it's effective)

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#57
post #45

Earlier quoted context omitted.

Unlike the big companies who do this, which often are merely impure scams on tax payer money a little more downstream.

Great, now we're defending embezzlement and fraud with public funds on HN, because we really really hate big business. A child caught doing something bad will cry "but my friends also did it!", is that the level of reasoning hackers want to be at?

What part of that said "defense?"

They can both be bad.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#58

Oh no, someone is profiting off of their work without proper attribution!?!?

This is an open weights model based on other open weights models. The dispute is that they released it with claims about having done some post training that improved the outputs. It was discovered that the model was not post trained like they claimed. The HF page now says it’s a merge of models, which wasn’t there before. They’re trying to claim they accidentally uploaded the wrong model to HF and that they’ll upload…

Thanks for the factual clarification. This is so important when everyone already has their trigger finger on politics. Not meaning that politics are irrelevant here, see sister comment by jobim.

But it's impossible to form a nuanced opinion when political association has a higher priority than the facts; which, again, don't look flattering for the implementers.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#60
post #45

Earlier quoted context omitted.

Unlike the big companies who do this, which often are merely impure scams on tax payer money a little more downstream.

Great, now we're defending embezzlement and fraud with public funds on HN, because we really really hate big business. A child caught doing something bad will cry "but my friends also did it!", is that the level of reasoning hackers want to be at?

That seems like a bad faith read to me. Nobody is defending it, just pointing out the irony / hypocrisy. Two things can be bad, and they can be related.
Post reply on HN