Live data from Hacker News

Groqchat

chat.groq.com

61–70 of 131 posts

Re: Groqchat

#61

I can't wait until LLMs are fast enough that a single response can actually be a whole tree of thought/review process before giving you an answer, yet is still fast enough to not even notice

I would bet a chunk of $$ that right before that point there will be a shift to bigger structures. Maybe MOE with individual tree of thought, or “town square consensus” or something.

Re: Groqchat

#63

In case, it's not blinding obvious to people. Groq are a hardware company that have built chips that are designed around the training and serving of machine models particularly targeted at LLMs. So the quality of the response isn't really what we're looking for here. We're looking for speed i.e. tokens per second. I actually have a final round interview with a subsidiary of Groq coming up and I'm very undecided as to…

tbh anyone can build fast hw for a single model, I’d audit their plan for a SW stack before joining. That said their arch is pretty unique so if they’re able to get these speeds it is pretty compelling

Re: Groqchat

#64
Great work! This is the fastest inference I have ever seen of any truly large language model (>=70b parameters).

Just FYI, you might want to fix autocorrect on iOS, your textbox seems to suppress it (at least for me).

Re: Groqchat

#66

Earlier quoted context omitted.

Doesn't the speed also depend on the number of people currently accessing it?

I assume it queues up requests if there are too many in flight to handle all of them at the same time, but it shows you the tokens per second (T/s) for every response, which is the number that matters (and presumably won't include the time spent in the queue).

You can be confused by why people are interested in the outputs, or you can think that the explanation is adequate. You can't do both at once.

Here, in fact, the explanation is not adequate. Let's analyze:

> Welcome Groq® Prompster! Are you ready to experience the world's fastest Large Language Model (LLM)?

"The world's fastest LLM? They must have made an LLM, I guess"

> We'd suggest asking about a piece of history, requesting a recipe for the holiday season, or copy and pasting in some text to be translated; "Make it French."

"Instructions, ok".

> This alpha demo lets you experience ultra-low latency performance using the foundational LLM, Llama 2, 70B created by Meta AI - running on Groq.

"So it's not their own LLM? It's just LLaMa? What's interesting about that? And what's Groq, the thing this is 'running on'? The website? I kind of guessed that, by virtue of being here."

> Like any AI demo, accuracy, correctness, or appropriateness cannot be guaranteed.

"Ok, sure".

Nowhere does it say "we've built new, very fast processing hardware for LLMs. Here's a demo a bog-standard LLM, LLaMa 70B, running on our hardware. Notice how fast it is".

Re: Groqchat

#67
post #43

Earlier quoted context omitted.

Thanks, impressive full-stack work. I'm sure this was named long before Musk decided to set 44B and change on fire but at first I confused it with Twitter's own LLM thing.

They addressed it on their blog https://groq.com/hey-elon-its-time-to-cease-de-grok/

can you trademark a verb ?

Re: Groqchat

#70

This isn't running on one chip. It's running on 128, or two racks worth of their kit. https://news.ycombinator.com/item?id=38739106 This doesn't mean much without comparing $ or watts of GPU equivalents

[deleted]
Post reply on HN