Live data from Hacker News

Groq runs Mixtral 8x7B-32k with 500 T/s

groq.com

331–340 of 482 posts

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#333

Earlier quoted context omitted.

Groq devices are really well set up for small-batch-size inference because of the use of SRAM. I'm not so convinced they have a Tok/sec/$ advantage at all, though, and especially at medium to large batch sizes which would be the groups who can afford to buy so much silicon. I assume given the architecture that Groq actually doesn't get any faster for batch sizes >1, and Nvidia cards do get meaningfully higher through…

I assume given the architecture that Groq actually doesn't get any faster for batch sizes >1 I guess if you don't have any extra junk you can pack more processing into the chip?

(Groq Employee) Yes! Determinism + Simplicity are superpowers for ALU and interconnect utilization rates. This system is powered by 14nm chips, and even the interconnects aren't best in class.

We're just that much better at squeezing tokens out of transistors and optic cables than GPUs are - and you can imagine the implications on Watt/Token.

Anyways.. wait until you see our 4nm. :)

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#334
post #156

Earlier quoted context omitted.

The speed part or the being swallowed part?

The speed part. We're not interested in being swallowed. The aim is to be bigger than Nvidia in three years :)

Is Sam going to give you some of his $7T to help with that?

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#335

Ignoring latency but not throughput, How does this compare in terms of Cost ( cards Acquisition cost and Power needed) with Nvidia GPU for inference?

We intend to be very competitive on cost, power, hardware, TCO, whatever it is. Custom-built silicon+hardware has the advantage in this space.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#336

Just a minor gripe the bullet option doesn't seem to be logical.. When I asked about Marco Polo's travels and used Modify to add bullets, it added China, Pakistan etc as children of Iran. And the same for other paragraphs.

(Groq Employee) Thanks for the feedback :) We're always improving that demo.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#337

Earlier quoted context omitted.

There's a difference between token throughput and latency. Token throughput is the token throughput of the whole GPU/system and latency is the token throughput for an individual user. Groq offers extremely low latency (aka extremely high token throughput per user) but we still don't have numbers on the token throughput of their entire system. Nvidia's metrics here on the other hand, show us the token throughput of th…

https://wow.groq.com/artificialanalysis-ai-llm-benchmark-dou... Seems to have it. Looks cost competitive but a lot faster.

People are using throughput and latency differently in different locations/contexts. Here they are referring to token throughput per user and first token/chunk latency. They don't mention the token throughput of the entire 576-chip system[0] that runs Llama 2 70b which would be the number we're looking for.

[0] https://news.ycombinator.com/item?id=38742581

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#338
post #323
post #72

Earlier quoted context omitted.

As a fellow scientist I concur with the approach of skepticism by default. Our chat app and API are available for everyone to experiment with and compare output quality with any other provider. I hope you are enjoying your time of having an empty calendar :)

Wait you have an API now??? Is it open, is there a waitlist? I’m on a plane but going to try to find that on the site. Absolutely loved your demo, been showing it around for a few months.

There is an API and there is a waitlist. Sign up at http://wow.groq.com/

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#339

Nice… a startup that has two “C” positions CEO and Chief Legal Officer… That sounds like a fun place to be

When you have a pile of hardware and silicon Intellectual Property, patents, etc, IMO it's pretty clever. However, I'm a Groq Engineer, and I'm mega-biased.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#340
post #139

I just want to say that this is one of the most impressive tech demos I’ve ever seen in my life, and I love that it’s truly an open demo that anyone can try without even signing up for an account or anything like that. It’s surreal to see the thing spitting out tokens at such a crazy rate when you’re used to watching them generate at one less than one fifth that speed. I’m surprised you guys haven’t been swallowed up…

Really glad you like it! We've been working hard on it.

Is this useful for training as well as running a model. Or is this approach specifically for running an already-trained model faster?
Post reply on HN