Live data from Hacker News

Groq runs Mixtral 8x7B-32k with 500 T/s

groq.com

431–440 of 482 posts

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#431

Earlier quoted context omitted.

Ask how much hardware is behind it.

All that matters is the cost. Their price is cheap, so the real question is whether they are subsidizing the cost to achieve that price or not.

> All that matters is the cost.

Not really, sustainability matters, if they are the only game in town, you want to know that game isn't going to end suddenly when their runway turns into a brick wall.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#432

Sorry, I'm a bit naïve about all of this. Why is this impressive? Can this result not be achieved by throwing more compute at the problem to speed up responses? Isn't the fact that there is a queue when under load just indicative that there's a trade-off between "# of request to process per unit of time" and "amount of compute to put into a response to respond quicker"? https://raw.githubusercontent.com/NVIDIA/Tensor…

I think NVidia is listing max throughput in terms of batching, so e.g. 50 tok/s for 10 different prompts at the same time. Groq LPUs definitely outerform an H100 in raw speed.

But fundamentally it's a system that only has 10x the speed for 500x the price, made by a company that runs a blockchain and is trying to heavily market what were intended to be crypto mining chips for LLM inference. It's really quite a funny coincidence that when someone amazed posts this weekly link there's an army of Groq engineers at the ready in the comments ready to say everything and anything.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#433
post #318

Earlier quoted context omitted.

This company is older than Elon's

ah ok cool, why the downvotes? did I offend more than one person with my ignorance? why did Elon name his Grok?

I don't know why you got downvoted, but "grok" is the Martian word for "understand deeply" from Robert Heinlein's "Stranger in a Strange Land".

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#435

Earlier quoted context omitted.

I presume that's because it's a custom asic not yet in mass production? If they can get costs down and put more dies into each card then it'll be business/consumer friendly. Let's see if they can scale production. Also, where tf is the next coral chip, alphabet been slacking hard.

I think Coral has been taken to the wooden shed out back. Nothing new out of them for years sadly

Yeah. And it's a real shame bc even before LLMs got big I was thinking, couple generations down the line and coral would be great for some home automation/edge AI stuff.

Fortunately LLMs and hard work of clever peeps running em on commodity hardware are starting to make this possible anyway.

Because Google Home/Assistant just seems to keep getting dumber and dumber...

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#436
post #423

Earlier quoted context omitted.

Are there voice responses in the demo? I couldn't find em?

Here's a live demo of CNN of Groq plugged into a voice API https://www.youtube.com/watch?v=pRUddK6sxDg&t=235s

Thanks, that's pretty impressive. I suppose with blazing fast token generation now things like diarisation and the actual model are holding us back.

Once it flawlessly understands when it is being spoken to/if it should speak based on the topic at hand (like we do) then it'll be amazing.

I wonder if ML models can feel that feeling of wanting to say something so bad but having to wait for someone else to stop talking first ha ha.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#438
How is the Token/second calculated? I ask it a simple prompt and the model generated a 150 word (about 300 tokens?) answer in 17 seconds, then mentioning the speed of 408T/s.

Also, I guess this demo would feel real time if you could stream the outputs to the UI? Can this be done in your current setup?

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#439

Earlier quoted context omitted.

It doesn't matter if you overheard it at a bar or if you're just some HN commenter posting completely incorrect legal advice; the law prohibits trading on material nonpublic information. I would pay a lot to see you try your ridiculous legal hokey-pokey on how to define an "insider."

If you did hear it in a bar, could you tweet it out before your trade, so the information is made public?

I really doubt that you can make yourself something public so that you can later act on it.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#440
post #345

Earlier quoted context omitted.

Is this useful for training as well as running a model. Or is this approach specifically for running an already-trained model faster?

Currently graphics processors work well for training. Language processors (LPUs) excel at inference.

Did you custom build those Language processors for this task? Or did you repurpose something already existing? I have never heard anyone use ‘Language processor’ before.
Post reply on HN