Live data from Hacker News

Mistral-8x7B-Chat

huggingface.co

21–30 of 75 posts

Re: Mistral-8x7B-Chat

#21
post #13

This model is better by many other contenders, but still far from GPT4. "what famous brands are there which change one letter from a common word to make a non-existent, but a catchy name, such as "musiq" instead of "music".. etc?" There are several brands that have played with words by changing a letter or adding a letter to create a new and memorable name. Here are a few examples: Qatar Airways - This airline's name…

LLMs work on tokens where characters are hidden away. They'd have to be explicitly trained on spelling each token out into single letter tokens and as they are bad at information symmetry - from single letter tokens back onto tokens as well. I don't think anybody does this so they're left with what's in training data only. Otherwise they don't have chance to reconstruct this information as tokens could map to any equ…

I thought so, too. But then I asked it to define fake words that were portmanteaus I made up. Believe me, my understanding of BERT and discriminant models aligned perfectly with what you're saying. But testing out the theory that it can break down and make meaning of fake words with accurate depictions of what words I'm combining proved me wrong. Generative models must work differently than you and I thought.

Re: Mistral-8x7B-Chat

#22
post #9

Earlier quoted context omitted.

what a sick project to be able to attract a billionaire programmer [0] and c royalty. [0]: https://github.com/ggerganov/llama.cpp/issues/4216#issuecomm...

For those out of the loop, who are the billionaire programmer and C royalty people in this link?

Tobi is the founder of Shopify.

Re: Mistral-8x7B-Chat

#23

There’s probably a better place to ask this highly specific technical question, but I’m avoiding Reddit these days so just throwing it out I guess. I’ve been trying to run these in a container but it’s verrrry slow, I believe, because of the lack of gpu help. All the instructions I find are for nvidia gpus and my server is a qnap tvs-473e with an embedded amd cpu/gpu (I know, I know). The only good news is that I’ve…

> my server is a qnap tvs-473e with an embedded amd cpu/gpu

That's your problem. I googled and it looks like one of these all-in-one appliances like a drobo or whatever's popular these days. That's not a server. (At least, I wouldn't call it a server. It's an all-in-one appliance, or toy, depending on perspective) And yegods, that price...

Spend $500, get an actual computer, not some priced up appliance, and you'll have a much better time. Regardless of if you spend it on more CPU or more GPU. You can get a used computer off ebay for $100 and shove a $400 graphics card in it. Or maybe get a ryzen 7 7700x, I'm looking at a mobo+cpu combo with that for $500 right now.

Finally, to make sure this response does contain a answer to what you asked: ;-)

if you can run this stuff in a container on your appliance already, but it's very slow, congrats! I'd call that a win. I looked up the chip, the RX-421BD, it's of similar power as an Athelon circa 2017. I think my router might have more compute power. You _do_ have those 512 shader cores, given effort, you could try and get them to do something useful. But I wouldn't assume it's possible (well, maybe you don't mind writing your own shaders ;-)). Just because the chip has "some gpu" doesn't mean it has "the right kind of gpu you'd need to hijack for lots of matrix multiplies, without writing the assembly yourself".

Sorry this isn't more helpful, but it's the truth.

Re: Mistral-8x7B-Chat

#24

There’s probably a better place to ask this highly specific technical question, but I’m avoiding Reddit these days so just throwing it out I guess. I’ve been trying to run these in a container but it’s verrrry slow, I believe, because of the lack of gpu help. All the instructions I find are for nvidia gpus and my server is a qnap tvs-473e with an embedded amd cpu/gpu (I know, I know). The only good news is that I’ve…

> qnap tvs-473e Specs say this runs an AMD RX-421BD. This is a 2015 AMD CPU with 2 bulldozer cores and a tiny IGP. ...To be blunt, you would be much better off running LLMs on your phone. Even an older phone. Or literally whatever device you are reading HN on. But if you insist , the runtime you want in MLC-LLM's Vulkan runtime.

This. Sibling llama.cpp comment is standard "I know llama.cpp, I assume that's 80% of the universe instead of .8%, and I assume that's all anyone needs. So I know just enough to be dangerous with ppl looking for advice".

You'll see it over and over again when you're looking for help, be careful, it's 100% a blind alley in your case. It's very likely you'll be disappointed by MLC as well, simultaneously it's your only real option. You definitely won't hit 1 tkn/sec, and honestly, id bet 0.1 tkn / sec

Re: Mistral-8x7B-Chat

#25
post #13

This model is better by many other contenders, but still far from GPT4. "what famous brands are there which change one letter from a common word to make a non-existent, but a catchy name, such as "musiq" instead of "music".. etc?" There are several brands that have played with words by changing a letter or adding a letter to create a new and memorable name. Here are a few examples: Qatar Airways - This airline's name…

But both are completely wrong! And technically the Google example is closer to correct than any others.

The Yi 34B eBay and Kodak examples are both (wrong but) very interesting because it does seem to get the idea of changing one letter.

Of GPT4 examples, the Qatar example (replacing "Q" with "Q" !?) is the only one that is internally consistent. The Pinterest and Tumblr examples are wrong in very odd ways in that the explanation doesn't match the spelling.

Re: Mistral-8x7B-Chat

#26

There’s probably a better place to ask this highly specific technical question, but I’m avoiding Reddit these days so just throwing it out I guess. I’ve been trying to run these in a container but it’s verrrry slow, I believe, because of the lack of gpu help. All the instructions I find are for nvidia gpus and my server is a qnap tvs-473e with an embedded amd cpu/gpu (I know, I know). The only good news is that I’ve…

You'll want to try llama.cpp [1]. The set of models that it can support is expanding [2]. Folks have also written services [3] that wrap around it. [1] https://github.com/ggerganov/llama.cpp [2] https://huggingface.co/TheBloke [3] https://github.com/abetlen/llama-cpp-python

Thanks! I was just following the thread about their recent addition of the OpenCl support and was on the verge of trying it out last weekend. I’ll definitely continue once I’m home again!

Re: Mistral-8x7B-Chat

#27

Man this LLM stuff gets released faster than I can keep up. Is there a centralized list somewhere that tests "use this for x purpose, use that for y?"

> Is there a centralized list somewhere that tests "use this for x purpose, use that for y?"

Yeah, "don't use these models for production, use OpenAI for production, ignore Claude/Gemini/etc.".

Re: Mistral-8x7B-Chat

#28
post #13

This model is better by many other contenders, but still far from GPT4. "what famous brands are there which change one letter from a common word to make a non-existent, but a catchy name, such as "musiq" instead of "music".. etc?" There are several brands that have played with words by changing a letter or adding a letter to create a new and memorable name. Here are a few examples: Qatar Airways - This airline's name…

Are you comparing a 8x7b model with GPT-4? Come on...

Re: Mistral-8x7B-Chat

#29

There’s probably a better place to ask this highly specific technical question, but I’m avoiding Reddit these days so just throwing it out I guess. I’ve been trying to run these in a container but it’s verrrry slow, I believe, because of the lack of gpu help. All the instructions I find are for nvidia gpus and my server is a qnap tvs-473e with an embedded amd cpu/gpu (I know, I know). The only good news is that I’ve…

> qnap tvs-473e Specs say this runs an AMD RX-421BD. This is a 2015 AMD CPU with 2 bulldozer cores and a tiny IGP. ...To be blunt, you would be much better off running LLMs on your phone. Even an older phone. Or literally whatever device you are reading HN on. But if you insist , the runtime you want in MLC-LLM's Vulkan runtime.

Thanks, I’ll look into it! Especially if the llama.cpp route is a dud, like the other response says it will be. My little qnap clunker handles all the self hosting stuff I throw at it, but I won’t be surprised if it simply has met its match

Re: Mistral-8x7B-Chat

#30
Somewhere between shiny Google releases and Mistral's magnet link tweet, there's gotta be a sweet spot where you release the model but also have enough decency to tell people how to use it optimally. Mistral, if you're reading this, I'm talking about you.
Post reply on HN