Live data from Hacker News

Groq runs Mixtral 8x7B-32k with 500 T/s

groq.com

411–420 of 482 posts

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#412

Earlier quoted context omitted.

there are providers out there offering for $0 per million tokens, that doesn't mean it is sustainable and won't disappear as soon as the VC well runs dry. Am not saying this is the case for Groq, but in general you probably should care if you want to build something serious on top of anything.

(Groq Employee) Agreed, one should care, and especially since this particular service is very differentiated by its speed and has no competitors. That being said, until there's another option at anywhere that speed.. That point is moot, isn't it :) For now, Groq is the only option that can let you build an UX with near-instant response times. Or a live agents that help with a human-to-human interaction. I could go on…

Why go so fast? Aren't Nvidias products fast enough from a TPS perspective?

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#413

Earlier quoted context omitted.

It doesn't matter if you overheard it at a bar or if you're just some HN commenter posting completely incorrect legal advice; the law prohibits trading on material nonpublic information. I would pay a lot to see you try your ridiculous legal hokey-pokey on how to define an "insider."

You're an idiot. https://www.kiplinger.com/article/investing/t052-c008-s001-w... Case #1.

Unless you earn enough money to retain good lawyers and are prepared to get into complicated legal troubles, getting sued isn't a great outcome even if you win.

The prudent thing to do is to stay away from anything that might make you become a target of investigation, unless the gains outweigh the risk by a significant margin.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#414

Earlier quoted context omitted.

They’re for sale on Mouser for $20625 each https://www.mouser.com/ProductDetail/BittWare/RS-GQ-GC1-0109... At that price 568 chips would be $11.7M

I presume that's because it's a custom asic not yet in mass production? If they can get costs down and put more dies into each card then it'll be business/consumer friendly. Let's see if they can scale production. Also, where tf is the next coral chip, alphabet been slacking hard.

I think Coral has been taken to the wooden shed out back. Nothing new out of them for years sadly

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#415

The main problem with the Groq LPUs is, they don't have any HBM on them at all. Just a miniscule (230 MiB) [0] amount of ultra-fast SRAM (20x faster than HBM3, just to be clear). Which means you need ~256 LPUs (4 full server racks of compute, each unit on the rack contains 8x LPUs and there are 8x of those units on a single rack) just to serve a single model [1] where as you can get a single H200 (1/256 of the server…

Groq Engineer here, I'm not seeing why being able to scale compute outside of a single card/node is somehow a problem. My preferred analogy is to a car factory: Yes, you could build a car with say only one or two drills, but a modern automated factory has hundreds of drills! With a single drill, you could probably build all sorts of cars, but a factory assembly line is only able to make specific cars in that configur…

>Show me a 30b+ parameter model doing RAG as part of a conversation with voice responses in less than a second, running on Nvidia.

Is your version of that on a different page from this chat bot?

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#417

Earlier quoted context omitted.

Groq states in this article [0] that they used 576 chips to achieve these results, and continuing with your analysis, you also need to factor in that for each additional user you want to have requires a separate KV cache, which can add multiple more gigabytes per user. My professional independent observer opinion (not based on my 2 years of working at Groq) would have me assume that their COGS to achieve these perfor…

What happened to Rex? Did it hit production or get abandoned? It was also on my list of things to consider modifying for an AI accelerator. :)

Long story, but technically REX is still around but has not been able to continue to develop due to lack of funding and my cofounder and I needing to pay bills. We produced initial test silicon, but due to us having very little money after silicon bringup, most of our conversations turned to acquihire discussions.

There should be a podcast release (https://microarch.club/) in the near future that covers REX's history and a lot of lessons learned.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#419

I always ask LLMs this: > If I initially set a timer for 45 minutes but decided to make the total timer time 60 minutes when there's 5 minutes left in the initial 45, how much should I add to make it 60? And they never get it correct.

gpt4 first go:

If you initially set a timer for 45 minutes and there are 5 minutes left, that means 40 minutes have already passed. To make the total timer time 60 minutes, you need to add an additional 20 minutes. This will give you a total of 60 minutes when combined with the initial 40 minutes that have already passed.

Re: Groq runs Mixtral 8x7B-32k with 500 T/s

#420
post #419

I always ask LLMs this: > If I initially set a timer for 45 minutes but decided to make the total timer time 60 minutes when there's 5 minutes left in the initial 45, how much should I add to make it 60? And they never get it correct.

gpt4 first go: If you initially set a timer for 45 minutes and there are 5 minutes left, that means 40 minutes have already passed. To make the total timer time 60 minutes, you need to add an additional 20 minutes. This will give you a total of 60 minutes when combined with the initial 40 minutes that have already passed.

Bard/Gemini gets it wrong the same way too. Interestingly, if I tell either GPT-4 or Gemini the right answer, they figure it out.
Post reply on HN