StableLM Zephyr 3B
11–20 of 42 posts
Re: StableLM Zephyr 3B
#12How fast are these small models on a 4090, is it like 100ms ? 500ms ?
Re: StableLM Zephyr 3B
#13"This model is being released under a non-commercial license that permits non-commercial use." I'm very interested in high quality 3B models, but it's hard to get excited about this given the increasing array of commercially usable models.
Re: StableLM Zephyr 3B
#14Re: StableLM Zephyr 3B
#15"This model is being released under a non-commercial license that permits non-commercial use." I'm very interested in high quality 3B models, but it's hard to get excited about this given the increasing array of commercially usable models.
It will be included under our membership next week which starts at $1 a month after grant ($20 base)
Re: StableLM Zephyr 3B
#16Re: StableLM Zephyr 3B
#17I might be missing it but do they say the number of training tokens that was used to train this?
This would help with efforts like TinyLlama in trying to figure out how well the scaling works with training tokens vs parameter size and challenging the chinchilla model.
Re: StableLM Zephyr 3B
#18Am I reading it right that performance was roughly comparable with GPT-3.5? How is this even possible?
Zephyr-7B-B still beats it in most benchmarks but it's close.
This model is almost Zephyr-7B-B performance at 3B size which is a lot better for inference requirements.
Re: StableLM Zephyr 3B
#19Am I reading it right that performance was roughly comparable with GPT-3.5? How is this even possible?
Re: StableLM Zephyr 3B
#20"This model is being released under a non-commercial license that permits non-commercial use." I'm very interested in high quality 3B models, but it's hard to get excited about this given the increasing array of commercially usable models.
Yeah, Replit is likely the best option out there for a 3B model size, right?