Live data from Hacker News

Mistral AI launches Mixtral-Next

chat.lmsys.org

21–30 of 55 posts

Re: Mistral AI launches Mixtral-Next

#21

AIExplained on youtube has guessed that Gemini 1.5 pro is taking Mistral’s accurate long content retrieval and Google just scaled it as much as they could. The Gemini 1.5 pro paper has a citation back to the last mistral paper in 2024.

And how does Mistral do "accurate long content retrieval"?

Re: Mistral AI launches Mixtral-Next

#23

Note that it's actually "Mistral Next" not "Mixtral Next" - so it isn't necessarily a MoE. For example, an early version of Mistral Medium (Miqu) was not a MoE but instead a Llama 70B model. I wonder how many parameters this one has

I know what they were going for with the Mixtral name but every time I come across it I wonder if they considered just how easily the two might be confused. It seems like a poor branding decision - what if some expected the Mixtral performance but accidentally uses a Mistral model? What if someone wants the low resource usage of e.g. Mistral 7B but tries out Mixtral 8x7B instead? It's especially hard when your colleagues aren't necessarily native English speakers.

There's got to be a better name for such a cool product. Maybe MistralX? MistMix?

Re: Mistral AI launches Mixtral-Next

#24
Mistral's process for releasing new models is extremely low-information. After getting very confused by this link I tried looking for a link that has any better information, and there just isn't one.

I thought Mixtral's release was weird when they just pasted a magnet link [0] into Twitter with no information, but at least people could download and analyze it so we got some reasonable third-party commentary in between that and the official announcement. With this one there's nothing at all to go on besides the name and the black box.

[0] https://news.ycombinator.com/item?id=38570537

Re: Mistral AI launches Mixtral-Next

#25

Earlier quoted context omitted.

I tried a bunch of my recent prompts to GPT-4 from daily use - this was often just slightly worse, sometimes slightly better. Fast too (tokens per second) while also not being overly wordy - very much appreciated that. Refusals are a bit "I am just a language model"-y which GPT-4 has gotten away from. Also it's more refuse-y if I broach something rudely (which again I've found GPT-4 to have become much better at.) Wa…

The clincher for me will be if the API to it is not so exorbitantly priced as GPT4, and if mistral can make using LoRAs economical.

[deleted]

Re: Mistral AI launches Mixtral-Next

#28

Mistral's process for releasing new models is extremely low-information. After getting very confused by this link I tried looking for a link that has any better information, and there just isn't one. I thought Mixtral's release was weird when they just pasted a magnet link [0] into Twitter with no information, but at least people could download and analyze it so we got some reasonable third-party commentary in betwee…

Company creates blackbox technology, and the company's communications are themselves like a blackbox... fitting

(I know that Mistral does a lot more stuff in the open than other companies, just couldn't resist the parallel between this and the blackbox limitations of LLMs in general)

Re: Mistral AI launches Mixtral-Next

#29

Mistral's process for releasing new models is extremely low-information. After getting very confused by this link I tried looking for a link that has any better information, and there just isn't one. I thought Mixtral's release was weird when they just pasted a magnet link [0] into Twitter with no information, but at least people could download and analyze it so we got some reasonable third-party commentary in betwee…

[deleted]
Post reply on HN