Live data from Hacker News

Mixtral 8x22B

mistral.ai

211–220 of 252 posts

Re: Mixtral 8x22B

#211

Curious to see how it performs against GPT-4. Mixtral8x22 beats CommandR+, which is at GPT-4-level in LMSYS' leaderboard.

LMSYS leaderboard is just one benchmark (that I think is fundamentally flawed). GPT-4 is clearly better.

Which alterative benchmarks do you recommend?

Re: Mixtral 8x22B

#212

Earlier quoted context omitted.

I really really like my Macbook Pro. But dammit, you can't upgrade the thing (Mac laptops aren't upgrade-able anymore). I got M1 Max in 2021 with 32GB of RAM. I did not anticipate needing more than 32GB for anything I'd be doing on it. Turns out, a couple of years later, I like to run local LLMs that max out my available memory.

I say 2021, but truth is the supply chain was so trash that year that it took almost a year to actually get delivered. I don't think I actually started using the thing until 2022.

I got downvoted for saying a true fact? that I ordered a the new M1 Max in 2021 and it took almost a year for me to actually get it? it true.

Re: Mixtral 8x22B

#213
post #209
post #94

Earlier quoted context omitted.

Wasn't there a paper yesterday that turned context evaluation linear (instead of quadratic) and made effectively unlimited context windows possible? Between that and 1.58b quantization I feel like we're overdue for an LLM revolution.

So far, people have come up with many alternatives for quadratic attention. Only recently have they proven their potential.

tons and tons of papers, most of them had some disadvantages. Can't have the cake and eat it too:

https://arxiv.org/html/2404.08801v1 Meta Megalodon

https://arxiv.org/html/2404.07143v1 Google Infini-Attention

https://arxiv.org/html/2402.13753v1 LongRoPE

and a ton more

Re: Mixtral 8x22B

#214
post #112

Earlier quoted context omitted.

You can't upgrade it? Edit: I haven't owned a laptop for years, probably could have surmised they'd be more user hostile nowadays.

You are getting downvoted because you vaguely suggested something negative about an Apple product, as is my comment below

People are absurd with their downvotes. I got downvoted for saying it took almost a year for my macbook to arrive once i ordered it. Its true. But its also true that supply chains were a wreck at the time. Apple wasn't the only tech gadget that took forever to arrive.

Re: Mixtral 8x22B

#215

It feels absolutely amazing to build an AI startup right now. It's as if your product automatically becomes cheaper, more reliable, and more scalable with each new major model release. - We first struggled with limited context windows [solved] - We had issues with consistent JSON ouput [solved] - We had rate limiting and performance issues for the large 3rd party models [solved] - Hosting our own OSS models for small…

> It's as if your product automatically becomes cheaper, more reliable, and more scalable with each new major model release. and so do your competitor's products.

Any business idea built almost exclusively on AI, without adding much value, is doomed from the start. AI is not good enough to make humans obsolete yet. But a well finetuned model can for sure augment what individual humans can do.

Re: Mixtral 8x22B

#216
post #188
post #14

"64K tokens context window" I do wish they had managed to extend it to at least 128K to match the capabilities of GPT-4 Turbo Maybe this limit will become a joke when looking back? Can you imagine reaching a trillion tokens context window in the future, as Sam speculated on Lex's podcast?

How useful is such a large input window when most of the middle isn't really used? I'm thinking mostly about coding. But wheb putting even say 20k tokens into the input, a good chunk doesn't seem to be "remembered" or used for the output

While you're 100% correct, they are working on ways to make the middle useful, such a "Needle in a Haystack" testing. When we say we wish for context length that large, I think it's implied we mean functionally. But you do make a really great point.

Re: Mixtral 8x22B

#217

I'm really excited about this model. Just need someone to quantize it to ~3 bits so it'll run on a 64GB MacBook Pro. I've gotten a lot of use from the 8x7b model. Paired with llamafile and it's just so good.

Shopping for a new mbp. Do you think going with more ram would be wise?

Unfortunately, yes. Get as much as you can stomach paying for.

Re: Mixtral 8x22B

#219

Earlier quoted context omitted.

Everything is soldered in these days. It's complete garbage. And most of the other vendors just copy Apple so even things like Lenovo have the same problems. The current state of laptops is such trash

These days with Apple Silicon, RAM is a part of the SoC. It's not even soldered on, it's a part of the chip. Although TBF, they also offer insane memory bandwidths.

Yes it’s almost like we got some benefit from iterating on these aspects of hardware design, as opposed to the typical HN grump characterisation of unbridled evil whenever a laptop isn’t exactly like how they were in 2006.

Re: Mixtral 8x22B

#220

Earlier quoted context omitted.

I say 2021, but truth is the supply chain was so trash that year that it took almost a year to actually get delivered. I don't think I actually started using the thing until 2022.

I got downvoted for saying a true fact? that I ordered a the new M1 Max in 2021 and it took almost a year for me to actually get it? it true.

Maybe they think that you aren’t meaningfully contributing to the conversation, which seems very, very likely.
Post reply on HN