Live data from Hacker News

Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

github.com

151–160 of 261 posts

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#151
post #32

Earlier quoted context omitted.

That's also something all the AI companies have been doing.

Lying about model capability is right now the lingua franca of the cloud AI business model, almost; they yes-and each other's lies because they are in a position of needing to generate interest, including going as far as needing to trigger regulatory capture. (It's not news to anyone who has worked in sales-led businesses that salespeople are prone to believing the claims of other salespeople, I guess).

> Lying about model capability is right now the lingua franca of the cloud AI business model

Lying about your lab's capabilities != Lying about model capability

Exaggerating the capabilities of a new model that you've actually trained in press bulletins can be called marketing. Merging two models and claiming that you trained a new model is plain lazy.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#153

Earlier quoted context omitted.

> A simple linear combination of every weight did not degrade the performance of the model, but enhanced it. Enhanced it on a couple benchmarks, supposedly. The game is to turn knobs until you get a benchmark run that shows an improvement, then ship it. There are a lot of fine tunes and chimera models on HuggingFace that are supposedly better at some specific test, but when you use them for anything else they're usua…

I don't think your last point is correct. Ablation, when done correctly, seems to increase the quality and typically also the performance too.

Abliterarion is a brute force technique that removes or silences parts of the model. It reduces performance because the abliterated elements aren’t perfectly isolated to censorship so other aspects suffer.

Many of the “uncensored” model providers also do some fine tuning on the models. Some of them target better benchmarks or other measures, but outside of the benchmarks and metrics they’re fine tuned for they are generally noticeably worse than the original model.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#154

Earlier quoted context omitted.

Great, now we're defending embezzlement and fraud with public funds on HN, because we really really hate big business. A child caught doing something bad will cry "but my friends also did it!", is that the level of reasoning hackers want to be at?

That seems like a bad faith read to me. Nobody is defending it, just pointing out the irony / hypocrisy. Two things can be bad, and they can be related.

You'd be surprised to hear then that I'm not the owner of any big company which embezzles tax payer money, and have never been involved in such.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#155

Earlier quoted context omitted.

That seems like a bad faith read to me. Nobody is defending it, just pointing out the irony / hypocrisy. Two things can be bad, and they can be related.

You'd be surprised to hear then that I'm not the owner of any big company which embezzles tax payer money, and have never been involved in such.

I don’t follow how that makes sense as a response to what I said?

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#156

I have no affiliation with them but here's what I think happened: 1. They claim the official model is based on Qwen 397B. It's likely they didn't disclose Nex Pro at all because Nex itself is based on the same base model (not saying they shouldn't). 2. The improvement would come from merging the weights PLUS on-policy distillation. The confusion is that the uploaded model didn't have the distillation at all. 3. It's…

> 2. The improvement would come from merging the weights PLUS on-policy distillation. The confusion is that the uploaded model didn't have the distillation at all.

They merged the base model with another lab’s fine tuned model. The improvements could have come from getting some of the fine tuned weights from the other model.

If they really had a better performing model that they “accidentally” forgot to upload, they could have uploaded the correct file by now.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#157

Earlier quoted context omitted.

The problem is that they claimed to have made a big achievement with their home grown post training, and they expected to receive a lot of praise for it. Then researchers looked at the weights and there is no post training at all. They are now attributing both models they merged, but their excuse for the lack of post training is to claim they accidentally uploaded the wrong files.

I’d believe they accidentally uploaded the wrong files if they uploaded the correct ones. To state that they accidentally uploaded something else and then not upload the correct version means they probably do not have anything and either hope people forget about this or they are scrambling to have something that is at least close to their original claim.

"Oops, we uploaded the wrong files" is the standard deflection every time people like this get caught.

Look up "Reflection 70B" drama.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#158

Earlier quoted context omitted.

You'd be surprised to hear then that I'm not the owner of any big company which embezzles tax payer money, and have never been involved in such.

I don’t follow how that makes sense as a response to what I said?

Why would I be a hypocrite for pointing out public fund embezzlement?

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#159

Earlier quoted context omitted.

I don’t follow how that makes sense as a response to what I said?

Why would I be a hypocrite for pointing out public fund embezzlement?

You’re not. The originally mentioned “big companies” are.

Re: Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model

#160

I have no affiliation with them but here's what I think happened: 1. They claim the official model is based on Qwen 397B. It's likely they didn't disclose Nex Pro at all because Nex itself is based on the same base model (not saying they shouldn't). 2. The improvement would come from merging the weights PLUS on-policy distillation. The confusion is that the uploaded model didn't have the distillation at all. 3. It's…

I'm honestly impressed that this even happened at all. "Rio de Janeiro's homegrown LLM" is probably the last headline I ever expected to read on HN.

Worth reminding everyone that Lua was also created in Rio, though admittedly at PUC rather than by the government.

Rio has a strong engineering talent pool, along with many other major capitals in Brazil

Post reply on HN