Live data from Hacker News

StableLM Zephyr 3B

stability.ai

11–20 of 42 posts

Re: StableLM Zephyr 3B

#13
post #6

"This model is being released under a non-commercial license that permits non-commercial use." I'm very interested in high quality 3B models, but it's hard to get excited about this given the increasing array of commercially usable models.

It will be included under our membership next week which starts at $1 a month after grant ($20 base)

Re: StableLM Zephyr 3B

#15
post #13
post #6

"This model is being released under a non-commercial license that permits non-commercial use." I'm very interested in high quality 3B models, but it's hard to get excited about this given the increasing array of commercially usable models.

It will be included under our membership next week which starts at $1 a month after grant ($20 base)

The parent comment was referring to free ($0 a month) models. With those, companies don't need to plan around the possibility that Stability AI hikes up the price afterwards.

Re: StableLM Zephyr 3B

#17
> Hardware: StableLM Zephyr 3B was trained on the Stability AI cluster across 8 nodes with 8 A100 80GBs GPUs for each nodes.

I might be missing it but do they say the number of training tokens that was used to train this?

This would help with efforts like TinyLlama in trying to figure out how well the scaling works with training tokens vs parameter size and challenging the chinchilla model.

Re: StableLM Zephyr 3B

#18
post #16

Am I reading it right that performance was roughly comparable with GPT-3.5? How is this even possible?

No it's not (according to their benchmarks).

Zephyr-7B-B still beats it in most benchmarks but it's close.

This model is almost Zephyr-7B-B performance at 3B size which is a lot better for inference requirements.

Re: StableLM Zephyr 3B

#19
post #16

Am I reading it right that performance was roughly comparable with GPT-3.5? How is this even possible?

By comparing on benchmarks that are either limited, or have data leaks, or in most cases just don't make sense in terms of usability - I've personally stopped looking at benchmarks to compare models. Personally, if I want to try a new model I hear a lot of chatter about, I use it for a few hours in my daily workflow. My baseline is GPT3.5 and GPT4, and I compare the models with them in terms of my day to day usage.

Re: StableLM Zephyr 3B

#20
post #6

"This model is being released under a non-commercial license that permits non-commercial use." I'm very interested in high quality 3B models, but it's hard to get excited about this given the increasing array of commercially usable models.

Yeah, Replit is likely the best option out there for a 3B model size, right?

Refact has a decent 1.6B model that I think is better

https://huggingface.co/smallcloudai/Refact-1_6B-fim

Post reply on HN