Live data from Hacker News

Mistral AI launches Mixtral-Next

chat.lmsys.org

41–50 of 55 posts

Re: Mistral AI launches Mixtral-Next

#41

Mistral's process for releasing new models is extremely low-information. After getting very confused by this link I tried looking for a link that has any better information, and there just isn't one. I thought Mixtral's release was weird when they just pasted a magnet link [0] into Twitter with no information, but at least people could download and analyze it so we got some reasonable third-party commentary in betwee…

> for releasing new models is extremely low-information

To be fair, this is not a release. This was the previous release https://mistral.ai/news/mixtral-of-experts/

It looks more like not trying very hard to hide things until release, rather than being a black box.

Re: Mistral AI launches Mixtral-Next

#42

Mistral's process for releasing new models is extremely low-information. After getting very confused by this link I tried looking for a link that has any better information, and there just isn't one. I thought Mixtral's release was weird when they just pasted a magnet link [0] into Twitter with no information, but at least people could download and analyze it so we got some reasonable third-party commentary in betwee…

> for releasing new models is extremely low-information To be fair, this is not a release. This was the previous release https://mistral.ai/news/mixtral-of-experts/ It looks more like not trying very hard to hide things until release, rather than being a black box.

If this were the first incident like this I would agree, but they very intentionally dropped the magnet link for Mixtral on Twitter with no further context. That leaves me wondering if this was also a weird on purpose thing rather than just them being casual.

Re: Mistral AI launches Mixtral-Next

#43

Earlier quoted context omitted.

> for releasing new models is extremely low-information To be fair, this is not a release. This was the previous release https://mistral.ai/news/mixtral-of-experts/ It looks more like not trying very hard to hide things until release, rather than being a black box.

If this were the first incident like this I would agree, but they very intentionally dropped the magnet link for Mixtral on Twitter with no further context. That leaves me wondering if this was also a weird on purpose thing rather than just them being casual.

Does it matter? You know that if you really want to play with things early, you may get an opportunity. And if you want to read more details, you'll get an announcement too. What's the problem with it being either on purpose or casual?

Re: Mistral AI launches Mixtral-Next

#44

Earlier quoted context omitted.

If this were the first incident like this I would agree, but they very intentionally dropped the magnet link for Mixtral on Twitter with no further context. That leaves me wondering if this was also a weird on purpose thing rather than just them being casual.

Does it matter? You know that if you really want to play with things early, you may get an opportunity. And if you want to read more details, you'll get an announcement too. What's the problem with it being either on purpose or casual?

It's an observation, not a complaint. It does leave very little to go on for an HN discussion, though, besides a meta conversation like this one.

Re: Mistral AI launches Mixtral-Next

#45

Earlier quoted context omitted.

I tried a bunch of my recent prompts to GPT-4 from daily use - this was often just slightly worse, sometimes slightly better. Fast too (tokens per second) while also not being overly wordy - very much appreciated that. Refusals are a bit "I am just a language model"-y which GPT-4 has gotten away from. Also it's more refuse-y if I broach something rudely (which again I've found GPT-4 to have become much better at.) Wa…

The clincher for me will be if the API to it is not so exorbitantly priced as GPT4, and if mistral can make using LoRAs economical.

Curious what your use case is that API tokens are actually significant expense at this time. I get it if you’re reselling a product but myself: abusively and naively messing around, I don’t see major prices as an end user. I sometimes hit the rate limit in ChatGPT, maybe like twice a week, but otherwise my monthly API bill in the last 6 months usually is in the dollars frame. Not tens or thousands of dollars, just dollars

Back when GPT4 came out mid last year my bills were slightly more dramatic (i am what you would probably call a semi heavy individual user) but they never surpassed $200 USD for a single month.

Re: Mistral AI launches Mixtral-Next

#47

Earlier quoted context omitted.

> for releasing new models is extremely low-information To be fair, this is not a release. This was the previous release https://mistral.ai/news/mixtral-of-experts/ It looks more like not trying very hard to hide things until release, rather than being a black box.

If this were the first incident like this I would agree, but they very intentionally dropped the magnet link for Mixtral on Twitter with no further context. That leaves me wondering if this was also a weird on purpose thing rather than just them being casual.

Well, what could they say? Given the lack of transparency on the data it well could be:

“We’ve trained LLaMA MoE on a lot of GPT4 data. And this it is not as good as GPT4. And this is our blob, so we can release it under any license. If someone is silly enough to use what this blob generates, this is not our problem.”

Re: Mistral AI launches Mixtral-Next

#49

wow, this might be the best LLM that i've used in terms of phrasing and presenting the answers.

I have the same impression. It is less verbose and doesn't beat around the bush (unlike GPT-3.5/4.5-Turbo), provides almost the same quality of code for the test cases I have, and has similar (GPT-4) or much better (GPT-3.5) spatial comprehension. It is at the GPT-3.5 level of math (read it as not good enough but better than anything else).

Re: Mistral AI launches Mixtral-Next

#50

Earlier quoted context omitted.

The clincher for me will be if the API to it is not so exorbitantly priced as GPT4, and if mistral can make using LoRAs economical.

Curious what your use case is that API tokens are actually significant expense at this time. I get it if you’re reselling a product but myself: abusively and naively messing around, I don’t see major prices as an end user. I sometimes hit the rate limit in ChatGPT, maybe like twice a week, but otherwise my monthly API bill in the last 6 months usually is in the dollars frame. Not tens or thousands of dollars, just do…

Oh, dude, you'll easily end up with $200 per hour if you are working on generating synthetic data...
Post reply on HN