Live data from Hacker News

Groq CEO: 'We No Longer Sell Hardware'

eetimes.com

111–120 of 152 posts

Re: Groq CEO: 'We No Longer Sell Hardware'

#111
post #78

Earlier quoted context omitted.

> 1.6 vs 1.8-2.2 seconds I believe certain companies would kill for 20% performance improvements on their main product.

"kill", .. why would anyone kill for a fraction of a second in this case? Informed folks know that LLM hosters aren't raking in the big bucks. They're selling dreams and aspirations, and those are what's driving the funding.

Google has used LMs in search for years (just not trendy LLMs), and search is famously optimized to the millisecond. Visa uses LMs to perform fraud detection every time someone makes a transaction, which is also quite latency sensitive. I'm guessing "informed folks" aren't so informed about the broader market.

OpenAI and Anthropic's APIs are obviously not latency-driven. Same with comparable LLM API resellers like Azure. Most people are likely not expecting tight latency SLOs there. That said, chat experiences (esp. voice ones) would probably be even more valuable if they could react in "human time" instead of with few seconds delay.

Integrating specialized hardware that can shave inference to fractions of a second seems like something that could be useful in a variety of latency-sensitive opportunities. Especially if this allows larger language models to be used where traditionally they were too slow.

Re: Groq CEO: 'We No Longer Sell Hardware'

#112
post #97

Earlier quoted context omitted.

I have lots of questions about how important latency is since you may be replacing many minutes or hours of a person’s time with undoubtedly a quicker response by any measure. This seems like a knee jerk reaction assuming latency is as important as it’s been with advertising. I’m not convinced latency matters as much as groqs material tries to claim it does.

When has latency ever not mattered? Let alone 'chat' use cases, but holding a reponse up for N*1.2 longer than it could holds all sorts of other resources up/down stream.

When it's already faster than I can absorb the response, which for me as an organic brain includes the normal token generation rate of the free tier of ChatGPT.

If I was using them to process far more text, e.g. summarise long documents, or if I was using it as an inline editing assistant, then I'd care more about the speed.

Re: Groq CEO: 'We No Longer Sell Hardware'

#113
post #53

Earlier quoted context omitted.

What open source model are you using when you hit groq? I just benchmarked some perf for some of my larger context window queries last week and groq's API took 1.6 seconds versus 1.8 to 2.2 for OpenAI GPT-3.5-turbo. So, it wasn't much faster. I almost emailed their support to see if I was doing something wrong. Would love to hear any details about your workload or the complexity of your queries.

What context did I miss that implies they are using an open source model?

If you go to GroqChat (which is like a demo app), they offer Gemma, Mistral, and LLaMa. These are all open-weights models.

Re: Groq CEO: 'We No Longer Sell Hardware'

#114
post #112
post #97

Earlier quoted context omitted.

When has latency ever not mattered? Let alone 'chat' use cases, but holding a reponse up for N*1.2 longer than it could holds all sorts of other resources up/down stream.

When it's already faster than I can absorb the response, which for me as an organic brain includes the normal token generation rate of the free tier of ChatGPT. If I was using them to process far more text, e.g. summarise long documents, or if I was using it as an inline editing assistant, then I'd care more about the speed.

> When it's already faster than I can absorb the response

Streaming a response from a chatbot is only one use-case of LLMs.

I would argue the most interesting applications do not fall into this category.

Re: Groq CEO: 'We No Longer Sell Hardware'

#115
post #43
post #22

Earlier quoted context omitted.

I'm not questioning the deployment strategy, I'm wondering why Saudi Aramco wants to access so much compute power that is highly specialized(?) for generative AI workloads. Or is it more general than that?

It’s about material discovery[0], Aramco built a 250b model to help them improve efficiency [1]. [0] https://deepmind.google/discover/blog/millions-of-new-materi... [1] https://www.aramco.com/en/news-media/speeches/2024/leap24--r...

What's the connection there? The second link says it's not about materials discovery at all but rather their GigaPOWERS model, which is a physics simulation used to optimize CO2 injection into their fields (i.e. optimizing recovery). POWERS is old, it was in development for decades already. Given that they don't plan to use Groq for LLMs but simply for its parallel computation abilities. I wonder to what extent this deal - if it goes through - will seriously drain Groq, actually, as POWERS would not be like the code Groq was designed to run and so much of their performance comes from the way they tightly optimize for very specific calculations.

Re: Groq CEO: 'We No Longer Sell Hardware'

#116

Earlier quoted context omitted.

> 1.6 vs 1.8-2.2 seconds I believe certain companies would kill for 20% performance improvements on their main product.

I have lots of questions about how important latency is since you may be replacing many minutes or hours of a person’s time with undoubtedly a quicker response by any measure. This seems like a knee jerk reaction assuming latency is as important as it’s been with advertising. I’m not convinced latency matters as much as groqs material tries to claim it does.

Depends on your application.

For example, if you're a game company and you want to use LLMs so your players can converse with nonplayer characters in natural language, replacing a multiple-choice conversation tree - you'd want that to be low latency, and you'd want it to be cheap.

Re: Groq CEO: 'We No Longer Sell Hardware'

#117
post #44

I don't understand why the comments are trash-talking Groq. They are the fastest LLM inference provider by a big margin. Why would they sell their hardware to any other company for any price? Keep it all for themselves and take over the market. 95% of my LLM requests go to Groq these days because it's 0.25 seconds round trip for a complete answer. In comparison, "Claude Instant" takes about 4 seconds. The other 5% of…

>> why the comments are trash-talking Groq

they probably bought NVDA stock :)

Re: Groq CEO: 'We No Longer Sell Hardware'

#118
post #112

Earlier quoted context omitted.

When it's already faster than I can absorb the response, which for me as an organic brain includes the normal token generation rate of the free tier of ChatGPT. If I was using them to process far more text, e.g. summarise long documents, or if I was using it as an inline editing assistant, then I'd care more about the speed.

> When it's already faster than I can absorb the response Streaming a response from a chatbot is only one use-case of LLMs. I would argue the most interesting applications do not fall into this category.

Number of different use cases (categories) I'd agree; I'm not so sure about use (volume)…

…not yet anyway. Fast moving area, lots of blue water outside the chat interface.

Re: Groq CEO: 'We No Longer Sell Hardware'

#119
post #43

Earlier quoted context omitted.

It’s about material discovery[0], Aramco built a 250b model to help them improve efficiency [1]. [0] https://deepmind.google/discover/blog/millions-of-new-materi... [1] https://www.aramco.com/en/news-media/speeches/2024/leap24--r...

What's the connection there? The second link says it's not about materials discovery at all but rather their GigaPOWERS model, which is a physics simulation used to optimize CO2 injection into their fields (i.e. optimizing recovery). POWERS is old, it was in development for decades already. Given that they don't plan to use Groq for LLMs but simply for its parallel computation abilities. I wonder to what extent this…

Hi there, I work for Groq.

Groq's system was designed to run arbitrary high performance numerical workloads. In the past it has been used for a variety of scientific computation tasks, including nuclear fusion and drug discovery.

https://www.alcf.anl.gov/news/argonne-deploys-new-groq-syste...

Re: Groq CEO: 'We No Longer Sell Hardware'

#120

Earlier quoted context omitted.

> 1.6 vs 1.8-2.2 seconds I believe certain companies would kill for 20% performance improvements on their main product.

I have lots of questions about how important latency is since you may be replacing many minutes or hours of a person’s time with undoubtedly a quicker response by any measure. This seems like a knee jerk reaction assuming latency is as important as it’s been with advertising. I’m not convinced latency matters as much as groqs material tries to claim it does.

Google won search in large part because of their latency. I stopped using local models because of latency. I switched from OpenAI to VertexAI because of latency (and availability)
Post reply on HN