Earlier quoted context omitted.
I am using the exact same model. Ryzen 5600G w/32GB and an Nvidia P40 w/24GB VRAM 20/33 layers offloaded to GPU, 4K context. Uses 25GB system RAM and all 24GB VRAM. 5-7 tokens per second.
and Groq does 485.08 T/s on mixtral 8x7B-32k I am not sure local models have any future other than POC/research. Depends on the cost of course.
Brave Leo now uses Mixtral 8x7B as default
181–184 of 184 posts
Re: Brave Leo now uses Mixtral 8x7B as default
#182Earlier quoted context omitted.
I just discovered Groq, which does 485.08 T/s on mixtral 8x7B-32k No idea on pricing but supposedly one can email to api@groq.com
I think you can try it online at chat.groq.com
Re: Brave Leo now uses Mixtral 8x7B as default
#183What are good API providers that serve mixtral? I know only octo ai which seems decent but will be good to know alternatives too
Re: Brave Leo now uses Mixtral 8x7B as default
#184It's nice using Brave because you have Chromium's better performance, without having to worry about Manifest V2 dying and taking adblocking down with it. I have uBlock Origin enabled, but it has barely caught anything that slipped past the browser filters.
If by performance you mean browser performance, you have more performance with Firefox nowadays. https://news.ycombinator.com/item?id=36770883