This is fascinating that it worked though. Can we just merge all the open weight models and get something better?
Everything is using Stable Diffusion as underlying model, then most of the usage is merged of checkpoints
71–80 of 261 posts
This is fascinating that it worked though. Can we just merge all the open weight models and get something better?
Everything is using Stable Diffusion as underlying model, then most of the usage is merged of checkpoints
Earlier quoted context omitted.
Attribution isn't the relevant part. Lying about your lab's capabilities is.
I do not see anyone lying. The model card says: > Post-trained from Qwen 3.5 397B The model card also says that they use an inference framework based on "SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs" by Shi et al.: https://arxiv.org/abs/2510.05069 So the sources seem properly attributed. They only claim that what they did to "Qwen 3.5 397B" has improved the LLM, including, a…
Earlier quoted context omitted.
This is a pure scam on tax payer money. But what else would be expected?
Unlike the big companies who do this, which often are merely impure scams on tax payer money a little more downstream.
Earlier quoted context omitted.
This is an open weights model based on other open weights models. The dispute is that they released it with claims about having done some post training that improved the outputs. It was discovered that the model was not post trained like they claimed. The HF page now says it’s a merge of models, which wasn’t there before. They’re trying to claim they accidentally uploaded the wrong model to HF and that they’ll upload…
How do they just splice two models together?
Earlier quoted context omitted.
Unlike the big companies who do this, which often are merely impure scams on tax payer money a little more downstream.
Great, now we're defending embezzlement and fraud with public funds on HN, because we really really hate big business. A child caught doing something bad will cry "but my friends also did it!", is that the level of reasoning hackers want to be at?
I might be missing something, but I don’t see anyone defending the the scams.
Not surprised
Can someone please explain or link to some information about how models are merged? Is this genuinely merging weights mathematically or some kind of distillation (presumably not if they’ve done zero training as the post suggests).
But yes, in general, merging refers to techniques that directly blend the weights of different models mathematically. It had a big moment of popularity ~2 years ago, with many so-called "Frankenmodels" popping up on leaderboards.
I tend to think of merging as belonging to the same general umbrella as things like "abliteration", or other techniques that surgically modify the weights of a model without a traditional training/tuning loop. Maxime Labonne is a great person to follow if you're interested in this general area.
Oh, I am so SHOCKED, so SHOCKED! /s
Explaining the joke: in Brazil, Rio de Janeiro is known as "Terra de bandido" (Gangster's Land).
Kinda like Chicago in the 20's or Naples and Palermo in the 90s.
> Every weight tensor in Rio is, to thousands of standard deviations, the same 0.6/0.4 blend of Nex and Qwen — across all 60 layers and every component of the network. Other finetunes cannot be explained as interpolations. I find it amazing how robust the current deep learning models are. A simple linear combination of every weight did not degrade the performance of the model, but enhanced it.