"Cheeseface just dropped the Blippy-7B model which is almost as good as the twinamp 34B model on the SwagCube benchmark when run locally as int8 and this shows that the gains made by the skibidi-70B model will probably filter down to the baseline Eras models in the next few weeks"
Mistral-8x7B-Chat
41–50 of 75 posts
Re: Mistral-8x7B-Chat
#42Man this LLM stuff gets released faster than I can keep up. Is there a centralized list somewhere that tests "use this for x purpose, use that for y?"
Re: Mistral-8x7B-Chat
#43Earlier quoted context omitted.
> my server is a qnap tvs-473e with an embedded amd cpu/gpu That's your problem. I googled and it looks like one of these all-in-one appliances like a drobo or whatever's popular these days. That's not a server. (At least, I wouldn't call it a server. It's an all-in-one appliance, or toy, depending on perspective) And yegods, that price... Spend $500, get an actual computer, not some priced up appliance, and you'll h…
Meh I bought this 5 years ago because there was a sale on 10tb hard drives and I thought “Why shouldn’t I become a data hoarder?” And now it runs homeassistant and frigate and MeTube and jellyfin and if it doesn’t work for ollama then I’ll probably just deal with it, lol.
Unrelated, cool, I hadn't heard of any of those 4 programs, I'm googling now and some look useful. Thanks! Possibly saving me some time in my next project...
Re: Mistral-8x7B-Chat
#44Earlier quoted context omitted.
what a sick project to be able to attract a billionaire programmer [0] and c royalty. [0]: https://github.com/ggerganov/llama.cpp/issues/4216#issuecomm...
For those out of the loop, who are the billionaire programmer and C royalty people in this link?
Re: Mistral-8x7B-Chat
#45Man this LLM stuff gets released faster than I can keep up. Is there a centralized list somewhere that tests "use this for x purpose, use that for y?"
Something more domain/use-specific would be great to have
Re: Mistral-8x7B-Chat
#46Every day: "Cheeseface just dropped the Blippy-7B model which is almost as good as the twinamp 34B model on the SwagCube benchmark when run locally as int8 and this shows that the gains made by the skibidi-70B model will probably filter down to the baseline Eras models in the next few weeks"
Re: Mistral-8x7B-Chat
#47Every day: "Cheeseface just dropped the Blippy-7B model which is almost as good as the twinamp 34B model on the SwagCube benchmark when run locally as int8 and this shows that the gains made by the skibidi-70B model will probably filter down to the baseline Eras models in the next few weeks"
Everyone just seems to be running experiments independently and then randomly drop some results, with basically no documentation. Sometimes the motivation is clearly VC money or paper exposure, but sometimes there is no apparent motivation... Or even no model card. Then when something works, others copy the script.
Not that I dont enjoy it. I find the sea of finetune generations fascinating.
Re: Mistral-8x7B-Chat
#48There’s probably a better place to ask this highly specific technical question, but I’m avoiding Reddit these days so just throwing it out I guess. I’ve been trying to run these in a container but it’s verrrry slow, I believe, because of the lack of gpu help. All the instructions I find are for nvidia gpus and my server is a qnap tvs-473e with an embedded amd cpu/gpu (I know, I know). The only good news is that I’ve…
Save yourself some time and buy a 4090 if you really want to be high tier (consumer range) you will have a much faster experience. Not only with text. Also stable diffusion etc
Re: Mistral-8x7B-Chat
#49This model is better by many other contenders, but still far from GPT4. "what famous brands are there which change one letter from a common word to make a non-existent, but a catchy name, such as "musiq" instead of "music".. etc?" There are several brands that have played with words by changing a letter or adding a letter to create a new and memorable name. Here are a few examples: Qatar Airways - This airline's name…
Are you comparing a 8x7b model with GPT-4? Come on...
Re: Mistral-8x7B-Chat
#50Earlier quoted context omitted.
But both are completely wrong! And technically the Google example is closer to correct than any others. The Yi 34B eBay and Kodak examples are both (wrong but) very interesting because it does seem to get the idea of changing one letter. Of GPT4 examples, the Qatar example (replacing "Q" with "Q" !?) is the only one that is internally consistent. The Pinterest and Tumblr examples are wrong in very odd ways in that th…
> Of GPT4 examples, the Qatar example (replacing "Q" with "Q" !?)... That appears to be from vizzah's testing of Mistral-8x7B-Chat rather than GPT4.
I thought the first was GPT4 and the second was mislabeled "Yi 34B Chat" when they meant "Mistral-8x7B-Chat"?
Otherwise why are they saying it is far from GPT4?