Brave Leo now uses Mixtral 8x7B as default
171–180 of 184 posts
Re: Brave Leo now uses Mixtral 8x7B as default
#172Earlier quoted context omitted.
I found the snap update notifications too annoying on Ubuntu, so I tried the ppa. But it the video plugin would crash. So back to Chrome for me.
Curious whether you've tried with the new (non-PPA) repo directly from Mozilla as of v122 [1]. I think the old PPA was also Mozilla, so I don't know what may have changed aside from being more publicly acknowledged. Might be worth a try? I don't have an Ubuntu VM at-hand but on Debian bookworm it installed fine, and (after tweaking one line in profiles.ini to point to my old ESR profile) it loaded and played Widevine…
Re: Brave Leo now uses Mixtral 8x7B as default
#173Does anyone know of a good chrome extension for AI page summarization? I tried a bunch of the top Google search hits, they work fine but are really bloated with superfluous features.
https://kagi.com/summarizer/index.html
https://help.kagi.com/kagi/api/summarizer.html
"Alternatively use Kagi Search browser extension (Chrome/Firefox) and you can use the most advanced Muriel model right from the extension."
Re: Brave Leo now uses Mixtral 8x7B as default
#174Earlier quoted context omitted.
I think this perspective comes from a lack of historical experience and hands-on experience overall. Nvidia more broadly has very impressive support for their GPUs. If you look at the support lifecycles for their Jetson hardware over time it's significantly worse. I encourage you to look at what support lifecycles have looked like, with the most "egregious" example being dropping of support for the Jetson Nano in fro…
I've used Jetson for a few projects as a hobby. Made an I2S Sodar array with a TX2. And some robotics projects with a Jetson AGX Xavier that I got to evaluate and then to work on. And a few both, professional and toy projects with versions of Jetson Nano kit and Xavier. But this was between 2017 and 2021 or so. About a year back, I took that very early version of AGX Xavier, that got released years ago. It wasn't eve…
Nice! I'm sorry if I seemed dismissive or even disrespectful, in my experience Jetson certainly has it's place (why I've been using them for years) but compared to "bring your distro, apt-get/.run Nvidia driver" they can be a serious shock for casual users. Then they see the performance...
> Orin Nano being that slow in [0], it looks like you've been trying it in Aug 2023. It maybe worth re-evaluating on the latest Jetpack, it had transitioned to CUDA 12.2, TensorRT 8.6, cuDNN 8.9. I would expect that recent popularity of ASR/TTS pipelines and LLMs was not completely missed by Jetpack maintainers (there are some tutorials here - https://www.jetson-ai-lab.com ). And recently released JetPack could be optimized a lot more for these workflows.
Interestingly WIS was recently bumped to CUDA 12.2, etc and the performance improvements were very marginal. WIS uses Ctranslate2 under the hood (same as faster-whisper) which offers among the best Whisper performance overall but doesn't benefit much from changes in these underlying libraries. In the end even if it somehow magically doubled performance (it doesn't and won't) that still places the latest generation ~$600 Jetson board 5x slower than an ancient yet still fully officially supported ~$100 GPU. Power and form-factor is an issue but for the voice assistant use case a Jetson board barely doing realtime with Whisper medium is unacceptable to me and the vast majority of our users. Our goal is sub one second voice command sessions from end of speech, to command execution, to TTS response and Jetson just can't provide that at any cost.
I'm glad there are community resources for Jetson platforms (which I'm aware of) but their existence underscores my point - you'll notice when perusing through there are often various hoops to jump through whereas anything else is basically "install driver, container toolkit, docker run" and it just works and works performantly. Basically CUDA x86_64 and discrete GPUs is native/expected/developed for, Jetson is almost always a bit of an edge case with rough edges (relatively) all over the place.
> And your project is very cool!
Thanks! In terms of your suggestion I certainly might but in the meantime, overall (based on my Jetson experience) as I said in that discussion I'm very reluctant to officially support the Jetson line with WIS. I'm almost certain it will blow-back on the project and cause support headaches for us while all the while providing a sub-optimal user experience.
Re: Brave Leo now uses Mixtral 8x7B as default
#175Earlier quoted context omitted.
This is a bit surprising to hear. Current Jetpack 6 is Ubuntu 22.04 - this is the current Ubuntu LTS release. There's nothing ancient about it, no? I'm pretty sure, if I go and check versions of CUDA, PyTorch, Tensorflow - it'd be also relatively recent. I'd suggest checking what examples are available, see what community is doing, see if what you need had already been tried - https://www.jetson-ai-lab.com From what…
I have a Jetson as well, and you are sorely mistaken. Just reading the doc pages everything seems nice and well, but Nvidia deprecates these little boards like no other. No support after you've bought the thing, and everything is kept frozen. (ie no new python, no new python dependencies, etc) What they aren't telling you is that specific sub-versions within each jetson/orin family board have differing support (ie no…
Software for Jetson boards should be viewed as firmware for these embedded/industrial devices. They get installed in a robot, MRI machine, etc with a specific bespoke application targeting what they came with and are never touched again -or- supported by some large commercial firm with the skillsets you describe.
I was as firm/absolute in my original reply as I was because anyone who thinks life with Jetson is similar to life with a discrete Nvidia GPU on x86_64 will be in for a huge shock and 95% of the time it will end up on their shelf in a year or two.
It's one thing when it's the latest random ARM SBC you bought for $50 with no vendor support, it's another thing entirely when you're spending > $600 (or $2000 as this started!!!!) on a Jetson.
Re: Brave Leo now uses Mixtral 8x7B as default
#176Earlier quoted context omitted.
While not being officially supported, rocm runs just fine on my 6700XT, i just have to set an env var(export HSA_OVERRIDE_GFX_VERSION=10.3.0)
Really? Does everything run? Even AI stuff? Do you have any links where I can read more about that?
Also I only tried it on linux, AFAIK windows is a lot more difficult to get running, if it works at all...
With llama-cpp, I successfully tried various LLMs(e.g. LLAMA 13B, Mixtral etc) with very solid performance. Even for models that don't fit in VRAM completely, performance can be surprisingly solid, as long as you compile with AVX extensions. (and your CPU supports those)
Stable Diffusion via ComfyUI also works very well. However, be aware of VRAM limitations with the larger SDXL variants, especially when running a heavy desktop environment.
Regarding setup guides/links, there isn't a good centralized resource sadly, so some tinkering is needed. Unlike some of those CUDA 1-click solutions, ROCm requires more manual setup, especially for the models only unofficially supported.
Here are a couple of links that might be helpful:
https://old.reddit.com/r/LocalLLaMA/comments/18ourt4/my_setu...
https://old.reddit.com/r/StableDiffusion/comments/ww436j/how...
In general the r/localllama & r/StableDiffusion subreddits are good places to search for info.
Re: Brave Leo now uses Mixtral 8x7B as default
#177Asked Mistral 8x7B for an essay on ham. It started telling me about Hamlet.
It must start from the beginning. Pig > piglet. Ham > Hamlet
Re: Brave Leo now uses Mixtral 8x7B as default
#178What are good API providers that serve mixtral? I know only octo ai which seems decent but will be good to know alternatives too
I just discovered Groq, which does 485.08 T/s on mixtral 8x7B-32k No idea on pricing but supposedly one can email to api@groq.com
Re: Brave Leo now uses Mixtral 8x7B as default
#179Much better than llama 70b.
Re: Brave Leo now uses Mixtral 8x7B as default
#180What are good API providers that serve mixtral? I know only octo ai which seems decent but will be good to know alternatives too
Together.ai seems to be the best, incredibly fast.
And try mixtral on chat.groq.com