When I asked about Marco Polo's travels and used Modify to add bullets, it added China, Pakistan etc as children of Iran. And the same for other paragraphs.
Groq runs Mixtral 8x7B-32k with 500 T/s
331–340 of 482 posts
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#332omg. i can’t believe how incredibly fast that is. and capable too. wow
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#333Earlier quoted context omitted.
Groq devices are really well set up for small-batch-size inference because of the use of SRAM. I'm not so convinced they have a Tok/sec/$ advantage at all, though, and especially at medium to large batch sizes which would be the groups who can afford to buy so much silicon. I assume given the architecture that Groq actually doesn't get any faster for batch sizes >1, and Nvidia cards do get meaningfully higher through…
I assume given the architecture that Groq actually doesn't get any faster for batch sizes >1 I guess if you don't have any extra junk you can pack more processing into the chip?
We're just that much better at squeezing tokens out of transistors and optic cables than GPUs are - and you can imagine the implications on Watt/Token.
Anyways.. wait until you see our 4nm. :)
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#334Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#335Ignoring latency but not throughput, How does this compare in terms of Cost ( cards Acquisition cost and Power needed) with Nvidia GPU for inference?
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#336Just a minor gripe the bullet option doesn't seem to be logical.. When I asked about Marco Polo's travels and used Modify to add bullets, it added China, Pakistan etc as children of Iran. And the same for other paragraphs.
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#337Earlier quoted context omitted.
There's a difference between token throughput and latency. Token throughput is the token throughput of the whole GPU/system and latency is the token throughput for an individual user. Groq offers extremely low latency (aka extremely high token throughput per user) but we still don't have numbers on the token throughput of their entire system. Nvidia's metrics here on the other hand, show us the token throughput of th…
https://wow.groq.com/artificialanalysis-ai-llm-benchmark-dou... Seems to have it. Looks cost competitive but a lot faster.
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#338Earlier quoted context omitted.
As a fellow scientist I concur with the approach of skepticism by default. Our chat app and API are available for everyone to experiment with and compare output quality with any other provider. I hope you are enjoying your time of having an empty calendar :)
Wait you have an API now??? Is it open, is there a waitlist? I’m on a plane but going to try to find that on the site. Absolutely loved your demo, been showing it around for a few months.
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#339Nice… a startup that has two “C” positions CEO and Chief Legal Officer… That sounds like a fun place to be
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#340I just want to say that this is one of the most impressive tech demos I’ve ever seen in my life, and I love that it’s truly an open demo that anyone can try without even signing up for an account or anything like that. It’s surreal to see the thing spitting out tokens at such a crazy rate when you’re used to watching them generate at one less than one fifth that speed. I’m surprised you guys haven’t been swallowed up…
Really glad you like it! We've been working hard on it.