Earlier quoted context omitted.
Brave"s support of Manifest V3 is totally dependent on Google and Chrome >Brave will support uBO and uMatrix so long as Google doesn’t remove underlying V2 code paths (which seem to be needed for Chrome for enterprise support, so should stay in the Chromium open source) https://twitter.com/BrendanEich/status/1534893414579249152
Yeah, but the Brave adblocker is built-in, it's not an extension.
Brave Leo now uses Mixtral 8x7B as default
21–30 of 184 posts
Re: Brave Leo now uses Mixtral 8x7B as default
#22If you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S... . The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto…
I prefer koboldcpp over llama.cpp. It’s easy to spilt between gpu/cpu on models larger than VRAM
Haven't really been able to justify upgrading to a 4090 or similar given I play so few new games these days.
Re: Brave Leo now uses Mixtral 8x7B as default
#23What are good API providers that serve mixtral? I know only octo ai which seems decent but will be good to know alternatives too
Re: Brave Leo now uses Mixtral 8x7B as default
#24What are good API providers that serve mixtral? I know only octo ai which seems decent but will be good to know alternatives too
we use openrouter but have had some inconsistency with speed. i hear fireworks is faster, swapping it out soon.
Re: Brave Leo now uses Mixtral 8x7B as default
#25Earlier quoted context omitted.
Why not normal RAM? Ryzen 5600 with 128GB DDR4 is perfectly fine to run mixtral 8bit, and costs less than $1000. GPUs are only needed if you can not wait 5 minutes for an answer, or for training.
Or if you want multiple sessions at the same time. Or if you want to do anything else with your machine while it's running. But realistically, 5 minutes is too long. It should be conversational, and for that you need at least 5 tokens per second. Which your Ryzen just can't do.
To be fair, a lot of people are using this for non-interactive work, like batching document analysis or offline processing of user generated content.
Re: Brave Leo now uses Mixtral 8x7B as default
#26What are good API providers that serve mixtral? I know only octo ai which seems decent but will be good to know alternatives too
Re: Brave Leo now uses Mixtral 8x7B as default
#27If you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S... . The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto…
Re: Brave Leo now uses Mixtral 8x7B as default
#28Re: Brave Leo now uses Mixtral 8x7B as default
#29If you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S... . The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto…
Dumb question, but how can a 32 bit number be converted to 2 bits and still be useful? It seems like magic.
Re: Brave Leo now uses Mixtral 8x7B as default
#30It's nice using Brave because you have Chromium's better performance, without having to worry about Manifest V2 dying and taking adblocking down with it. I have uBlock Origin enabled, but it has barely caught anything that slipped past the browser filters.
Is this submarine comment?
> A "submarine comment" on social media refers to a comment that is made on an old post or thread, long after the conversation has died down. This term derives from the idea of a submarine which remains submerged and out of sight for long periods before suddenly surfacing. In the context of social media, it's when someone delves deep into someone else's posts or timeline, finds an old post, and leaves a comment, bringing the old post back to attention. This can sometimes surprise the original poster and other participants, as the conversation was thought to have been concluded.
Which doesn’t make sense in this context