Live data from Hacker News

Eagle 7B: Soaring past Transformers

blog.rwkv.com

71–80 of 86 posts

Re: Eagle 7B: Soaring past Transformers

#71
post #8

This shows the model architecture, be it transformer, Mamba, SSM or RWKV - doesn't really matter when compared to the impact of the training set. We're spending too much time debating models when we should be talking about language data, a reservoir of human experience won at great sacrifice by humanity. And the same data when used to train humans creates modern capable people. Alone, without society and language, we…

Model architecture matters. RWKV takes significantly less energy than transformer models.

It’s better for the environment and much faster.

It’s not _only_ about performance

Re: Eagle 7B: Soaring past Transformers

#72

Sounds like this hallucinates pretty easily From Reddit : https://www.reddit.com/r/LocalLLaMA/comments/1ad0j63/rwkv5_1... User: Which is larger, a chicken egg or a cow egg? Assistant: To determine which is larger, a chicken egg or a cow egg, let's first look at their respective sizes and compare them. Chicken Egg: The average chicken egg size ranges from 2.5 to 3 inches (6 to 8 cm) in length and 1.5 to 2 inches (3.8…

I absolutely should not anthropomorphise LLMs, but I can't get rid of the feelling that "it" was writing this answer with a mischievous smirk, and was having a lot of fun in the process.

The future is weird.

Re: Eagle 7B: Soaring past Transformers

#73
post #11
post #8

This shows the model architecture, be it transformer, Mamba, SSM or RWKV - doesn't really matter when compared to the impact of the training set. We're spending too much time debating models when we should be talking about language data, a reservoir of human experience won at great sacrifice by humanity. And the same data when used to train humans creates modern capable people. Alone, without society and language, we…

Of course it is in the brain, the brain created and evolved the language as a very powerful tool. If intelligence was in the language then other animals would be as intelligent as us

Other animals don't have our advanced language. In fact it is the lack of language transmission that keeps them down. What I am arguing is that we're pretty limited individually, only together, and with plenty of time, do we get so smart.

LLMs learning from the same text and gaining human like capabilities shows just how much of intelligence is crystalized in culture. If it works without brains, then brains were not the essential ingredient.

Humans without culture would need 10,000 years or more to recover, and have to pay the same price as the first time around. Culture is smarter than us.

Re: Eagle 7B: Soaring past Transformers

#74
post #8

This shows the model architecture, be it transformer, Mamba, SSM or RWKV - doesn't really matter when compared to the impact of the training set. We're spending too much time debating models when we should be talking about language data, a reservoir of human experience won at great sacrifice by humanity. And the same data when used to train humans creates modern capable people. Alone, without society and language, we…

It's very much worth discussing architecture since, if Mamba or Based end up working as good as Transformers, a lot of current problems related to quadratic scaling are solved.

Re: Eagle 7B: Soaring past Transformers

#75
post #53

> [Mar 2024] An MoE model based on the v5 Eagle 2T model (note, approximate date) Hyped about this! This could strike a powerful balance between performance and reasonably retained low environmental/token cost impact. Would be cool with improved coverage of Scandinavian languages along with it, but I guess we'll see. And yeah, I think a true revolution will happen (or might already be) when we realize the value of tr…

Looking into the nordic pile maybe? There are some datasets

Re: Eagle 7B: Soaring past Transformers

#76
post #69
post #58

Earlier quoted context omitted.

I also tried the demo and I find it pretty much useless at most things even comparing it to a small 7b transformer model like mistral. From my albeit quick tests, what I found is that it knows clearly less things than mistral, it hallucinates much more, it does not follow instructions, has less reasoning capabilities and asking it to translate a Japanese text into English gave me a bad translated summary instead of t…

As written in the post, it is a base model with light instruction tuning, i.e. Llama2, not Llama2-chat. You should evaluate it as a base model. If you evaluate it as a chat model, of course it will perform horribly.

OK, I missed that. Thanks for the clarification.

Re: Eagle 7B: Soaring past Transformers

#77
post #8

This shows the model architecture, be it transformer, Mamba, SSM or RWKV - doesn't really matter when compared to the impact of the training set. We're spending too much time debating models when we should be talking about language data, a reservoir of human experience won at great sacrifice by humanity. And the same data when used to train humans creates modern capable people. Alone, without society and language, we…

> We're spending too much time debating models when we should be talking about language data

Models still matter a lot. There's arguably still an abilities gap between LLMs and general intelligence that can only be bridged by a new model.

Re: Eagle 7B: Soaring past Transformers

#79
post #24
post #8

This shows the model architecture, be it transformer, Mamba, SSM or RWKV - doesn't really matter when compared to the impact of the training set. We're spending too much time debating models when we should be talking about language data, a reservoir of human experience won at great sacrifice by humanity. And the same data when used to train humans creates modern capable people. Alone, without society and language, we…

I agree that on balance, we should spend more effort on data than modeling, but it is just not true that modeling doesn't matter. Transformer-2023 is different from Trnasformer-2020 and cumulative improvement is significant. https://arxiv.org/abs/2312.00752 did such benchmark. If choice between Transformer and RWKV doesn't seem to matter to you, the only reason is that while Transformer-2020 evolved to Transformer-20…

They describe Transformer++ as "A Transformer with an improved architecture, namely rotary positional encodings and SwiGLU MLP", no linear bias terms and RMSNorm instead of LayerNorm.

But modern transformers have many more tricks than that. Such as pre-norm, sparse better use of residual layers, sparse attention masks and so on.

Re: Eagle 7B: Soaring past Transformers

#80
post #66

>> An eagle, flying past a transformer-looking robot That hero image is a complete mess (e.g. look at the eagle's forward paw, the "transformer robot"'s right arm or the position of the eagle's left wing). Why is it that people put up such obviously messed-up images in their articles? Do they not see that level of detail, or do they just find it cool to have some "AI art" in their article, as a kind of an in-group co…

I mean, They're an organisation creating AI models, releasing an AI model. Them using AI images isn't hugely shocking to anyone with the capacity to reason. The picture looks cool at first glance, and that is really all that matters. It's just a cool bit of eye-candy leading into the article.

>> Them using AI images isn't hugely shocking to anyone with the capacity to reason.

Gee, thanks, that's so kind.

Post reply on HN