If you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S... . The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto…
> You would need multiple GPUs with shared memory if you wanted to offload the higher precision models to VRAM. Or just a powerful apple silicon machine? I've tried dolphin mixtral 4bit on a 36gb ram MacBook m3, and inference is super fast.
Brave Leo now uses Mixtral 8x7B as default
31–40 of 184 posts
Re: Brave Leo now uses Mixtral 8x7B as default
#32It's nice using Brave because you have Chromium's better performance, without having to worry about Manifest V2 dying and taking adblocking down with it. I have uBlock Origin enabled, but it has barely caught anything that slipped past the browser filters.
Re: Brave Leo now uses Mixtral 8x7B as default
#33Re: Brave Leo now uses Mixtral 8x7B as default
#34Earlier quoted context omitted.
> You would need multiple GPUs with shared memory if you wanted to offload the higher precision models to VRAM. Or just a powerful apple silicon machine? I've tried dolphin mixtral 4bit on a 36gb ram MacBook m3, and inference is super fast.
I can run 4bit on a beat up 1070 ti. GP talks about higher precision models
Re: Brave Leo now uses Mixtral 8x7B as default
#35If you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S... . The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto…
Dumb question, but how can a 32 bit number be converted to 2 bits and still be useful? It seems like magic.
Re: Brave Leo now uses Mixtral 8x7B as default
#36Earlier quoted context omitted.
Or if you want multiple sessions at the same time. Or if you want to do anything else with your machine while it's running. But realistically, 5 minutes is too long. It should be conversational, and for that you need at least 5 tokens per second. Which your Ryzen just can't do.
>It should be conversational, and for that you need at least 5 tokens per second. To be fair, a lot of people are using this for non-interactive work, like batching document analysis or offline processing of user generated content.
Re: Brave Leo now uses Mixtral 8x7B as default
#37If you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S... . The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto…
I prefer koboldcpp over llama.cpp. It’s easy to spilt between gpu/cpu on models larger than VRAM
Re: Brave Leo now uses Mixtral 8x7B as default
#38If you want to run Mixtral 8x7B locally you can use llama.cpp (including with any of the supporting libraries/interfaces such as text-generation-webui) with https://huggingface.co/TheBloke/Nous-Hermes-2-Mixtral-8x7B-S... . The smallest quantized version (2bit) needs 20GB of RAM (which can be offloaded onto the VRAM of a decent 4090 GPU). The 4bit quantized versions are the largest models that can just about fit onto…
Dumb question, but how can a 32 bit number be converted to 2 bits and still be useful? It seems like magic.
Re: Brave Leo now uses Mixtral 8x7B as default
#39Just checking: PDF summarization is not yet implemented, right?
Re: Brave Leo now uses Mixtral 8x7B as default
#40Earlier quoted context omitted.
Is this submarine comment?
What is the definition of a submarine comment? Google fails and ChatGPT says: > A "submarine comment" on social media refers to a comment that is made on an old post or thread, long after the conversation has died down. This term derives from the idea of a submarine which remains submerged and out of sight for long periods before suddenly surfacing. In the context of social media, it's when someone delves deep into s…