Live data from Hacker News

Groq runs Mixtral 8x7B-32k with 500 T/s

groq.com

461–470 of 482 posts

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#461

I just want to say that this is one of the most impressive tech demos I’ve ever seen in my life, and I love that it’s truly an open demo that anyone can try without even signing up for an account or anything like that. It’s surreal to see the thing spitting out tokens at such a crazy rate when you’re used to watching them generate at one less than one fifth that speed. I’m surprised you guys haven’t been swallowed up…

ok... why tho? genuinely ignorant and extremely curious.

what's the TFLOPS/$ and TFLOPS/W and how does it compare with Nvidia, AMD, TPU?

from quick Googling I feel like Groq has been making these sorts of claims since 2020 and yet people pay a huge premium for Nvidia and Groq doesn't seem to be giving them much of a run for their money.

of course if you run a much smaller model than ChatGPT on similar or more powerful hardware it might run much faster but that doesn't mean it's a breakthrough on most models or use cases where latency isn't the critical metric?

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#462

Earlier quoted context omitted.

Curiously, almost all of this video is mostly covered by computer architectures lit in the late 90's early 00's. At the time, I recall Tom Knight had done most of the analysis in this video, but I don't know if he ever published it. It was extrapolating into the distant future. To answer your questions: - Spatial processors are an insanely good fit for async logic - Matrix of processing engines are a moderately good…

Sweet, thanks! It seems like this research ecosystem was incredibly rich, but Moore's law was in full swing, and statically known workloads weren't useful at the compute scale of back then. So these specialized approach never stood a chance next to CPUS. Nowadays the ground is.. more fertile.

Lots of things were useful to compute.

The problem was

1) If you took 3 years longer to build a SIMD architecture than Intel to make a CPU, Intel would be 4x faster by the time you shipped.

2) If, as a customer, I was to code to your architecture, and it took me 3 more years to do that, by that point, Intel would be 16x faster

And any edge would be lost. The world was really fast-paced. Groq was founded in 2016. It's 2024. If it was still hayday of Moore's Law, you'd be competing with CPUs running 40x as fast as today's.

I'm not sure you'd be so competitive against a 160GHz processor, and I'm not sure I'd be interested knowing a 300+GHz was just around the corner.

Good ideas -- lots of them -- lived in academia, where people could prototype neat architectures on ancient processes, and benchmark themselves to CPUs of yesteryear from those processes.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#463

Earlier quoted context omitted.

ah ok cool, why the downvotes? did I offend more than one person with my ignorance? why did Elon name his Grok?

I don't know why you got downvoted, but "grok" is the Martian word for "understand deeply" from Robert Heinlein's "Stranger in a Strange Land".

Thanks for the reply. It’s ok I’m trying to keep my karma under 1000

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#464

Earlier quoted context omitted.

Distributed, shared memory machines used to do exactly that in HPC space. They were a NUMA alternative. It works if the processing plus high-speed interconnect are collectively faster than the request rate. The 8x setups with NVLink are kind of like that model. You may have meant that nobody has a stack that uses clustering or DSM with low-latency interconnects. If so, then that might be worth developing given prior…

> Distributed, shared memory machines used to do exactly that in HPC space. reformed HPC person here. Yes, but not latency optimised in the case here. HPC is normally designed for throughput. Accessing memory from outside your $locality is normally horrifically expensive, so only done when you can't avoid it. For most serving cases, you'd be much happier having a bunch of servers with a number of groqs in them, than…

It would probably be a cluster of thin nodes with GPU’s or low-cost accelerators over a low-latency interconnect. The DSM would be layered on top of that. The AI cluster would handle processing with security, etc done more by other components. They’re usually layered.

I agree it’s harder to manage with less, fine-grained security. People were posting Groq chips at $20k each, though. With that, we’re talking whether the management of it is worth it for installations costing six or more digits. That might be more justifiable if an alternative saves them a good chunk of six or more digits.

Their main advantage is a solution that’s ready to go :)

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#465
post #156

Earlier quoted context omitted.

The speed part or the being swallowed part?

The speed part. We're not interested in being swallowed. The aim is to be bigger than Nvidia in three years :)

Why wouldn't NVIDIA release their own LPU?

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#466
post #423

Earlier quoted context omitted.

Are there voice responses in the demo? I couldn't find em?

Here's a live demo of CNN of Groq plugged into a voice API https://www.youtube.com/watch?v=pRUddK6sxDg&t=235s

Wow! Absolutely astounding!

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#467

I just want to say that this is one of the most impressive tech demos I’ve ever seen in my life, and I love that it’s truly an open demo that anyone can try without even signing up for an account or anything like that. It’s surreal to see the thing spitting out tokens at such a crazy rate when you’re used to watching them generate at one less than one fifth that speed. I’m surprised you guys haven’t been swallowed up…

If I understand correctly, each chip has 200 MB of RAM, so it takes racks to run a single LLM.

That does not sound like progress to me.

We need single PCIe boards with dozens or hundreds of GB of RAM and processors that handle it well.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#468
post #69

Earlier quoted context omitted.

I asked it to come up with name ideas for a company and it hallucinated them successfully :) I think the trick is to know what prompts will likely to yield results that are not likely to be hallucinated. In other contexts it's a feature.

A bit of a softball don't you think? The initial message suggests "Are you ready to experience the world's fastest Large Language Model (LLM)? We'd suggest asking about a piece of history" So I did.

My comment was about generic experience with LLM-s. Obviously your experience can differ from this.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#469

The main problem with the Groq LPUs is, they don't have any HBM on them at all. Just a miniscule (230 MiB) [0] amount of ultra-fast SRAM (20x faster than HBM3, just to be clear). Which means you need ~256 LPUs (4 full server racks of compute, each unit on the rack contains 8x LPUs and there are 8x of those units on a single rack) just to serve a single model [1] where as you can get a single H200 (1/256 of the server…

Groq Engineer here, I'm not seeing why being able to scale compute outside of a single card/node is somehow a problem. My preferred analogy is to a car factory: Yes, you could build a car with say only one or two drills, but a modern automated factory has hundreds of drills! With a single drill, you could probably build all sorts of cars, but a factory assembly line is only able to make specific cars in that configur…

Hi Matanyal, we worked with groq for a project last year. would you be open to connect on LinkedIn? :)

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#470

Earlier quoted context omitted.

there are providers out there offering for $0 per million tokens, that doesn't mean it is sustainable and won't disappear as soon as the VC well runs dry. Am not saying this is the case for Groq, but in general you probably should care if you want to build something serious on top of anything.

(Groq Employee) Agreed, one should care, and especially since this particular service is very differentiated by its speed and has no competitors. That being said, until there's another option at anywhere that speed.. That point is moot, isn't it :) For now, Groq is the only option that can let you build an UX with near-instant response times. Or a live agents that help with a human-to-human interaction. I could go on…

Hi foundval, can we connect on Linkedin please? :
Post reply on HN