Live data from Hacker News

Zebra-Llama – Towards efficient hybrid models

arxiv.org

1–10 of 66 posts

Re: Zebra-Llama – Towards efficient hybrid models

#3
> Zebra-Llama achieves Transformer-level accuracy with near-SSM efficiency using only 7–11B training tokens (compared to trillions of tokens required for pre-training) and an 8B teacher. Moreover, Zebra-Llama dramatically reduces KV cache size—down to 3.9%, 2%, and 2.73% of the original for the 1B, 3B, and 8B variants, respectively—while preserving 100%, 100%, and 97% of average zero-shot performance on LM Harness tasks.

This is an extraordinary claim, is there a catch I’m missing? Am I misreading?

Re: Zebra-Llama – Towards efficient hybrid models

#4
Due to perverse incentives and the historical nature of models over-claiming accuracy, it's very hard to believe anything until it is open source and can be tested out

that being said, I do very much believe that computational efficiency of models is going to go up [correction] drastically over the coming months, which does pose interesting questions over nvidia's throne

*previously miswrote and said computational efficiency will go down

Re: Zebra-Llama – Towards efficient hybrid models

#5

Due to perverse incentives and the historical nature of models over-claiming accuracy, it's very hard to believe anything until it is open source and can be tested out that being said, I do very much believe that computational efficiency of models is going to go up [correction] drastically over the coming months, which does pose interesting questions over nvidia's throne *previously miswrote and said computational ef…

I think you mean computational efficiency will go _up_ in the future. To your last point: Jevons paradox might apply.

Re: Zebra-Llama – Towards efficient hybrid models

#6

Due to perverse incentives and the historical nature of models over-claiming accuracy, it's very hard to believe anything until it is open source and can be tested out that being said, I do very much believe that computational efficiency of models is going to go up [correction] drastically over the coming months, which does pose interesting questions over nvidia's throne *previously miswrote and said computational ef…

Like this?

https://huggingface.co/amd/Zebra-Llama-8B-8MLA-24Mamba-SFT

Re: Zebra-Llama – Towards efficient hybrid models

#7

Due to perverse incentives and the historical nature of models over-claiming accuracy, it's very hard to believe anything until it is open source and can be tested out that being said, I do very much believe that computational efficiency of models is going to go up [correction] drastically over the coming months, which does pose interesting questions over nvidia's throne *previously miswrote and said computational ef…

I think you mean computational efficiency will go _up_ in the future. To your last point: Jevons paradox might apply.

yup that's what I meant!, Jevon's paradox applies to resource usage in general and not towards a specific companies dominance

if computational efficiency goes up (thanks for the correction), and CPU inference becomes viable for most practical applications, GPUs (or accelerators) themselves may be unnecessary for most practical functions

Re: Zebra-Llama – Towards efficient hybrid models

#8

Due to perverse incentives and the historical nature of models over-claiming accuracy, it's very hard to believe anything until it is open source and can be tested out that being said, I do very much believe that computational efficiency of models is going to go up [correction] drastically over the coming months, which does pose interesting questions over nvidia's throne *previously miswrote and said computational ef…

Like this? https://huggingface.co/amd/Zebra-Llama-8B-8MLA-24Mamba-SFT

yes!, thanks for the link!

Re: Zebra-Llama – Towards efficient hybrid models

#9

Earlier quoted context omitted.

I think you mean computational efficiency will go _up_ in the future. To your last point: Jevons paradox might apply.

yup that's what I meant!, Jevon's paradox applies to resource usage in general and not towards a specific companies dominance if computational efficiency goes up (thanks for the correction), and CPU inference becomes viable for most practical applications, GPUs (or accelerators) themselves may be unnecessary for most practical functions

Discrete GPUs still have an advantage in memory bandwidth. Though this might push platforms like laptops towards higher bandwidths, which would be nice.
Post reply on HN