I can't wait until LLMs are fast enough that a single response can actually be a whole tree of thought/review process before giving you an answer, yet is still fast enough to not even notice
Groqchat
61–70 of 131 posts
Re: Groqchat
#62Re: Groqchat
#63In case, it's not blinding obvious to people. Groq are a hardware company that have built chips that are designed around the training and serving of machine models particularly targeted at LLMs. So the quality of the response isn't really what we're looking for here. We're looking for speed i.e. tokens per second. I actually have a final round interview with a subsidiary of Groq coming up and I'm very undecided as to…
Re: Groqchat
#64Just FYI, you might want to fix autocorrect on iOS, your textbox seems to suppress it (at least for me).
Re: Groqchat
#65[flagged]
Re: Groqchat
#66Earlier quoted context omitted.
Doesn't the speed also depend on the number of people currently accessing it?
I assume it queues up requests if there are too many in flight to handle all of them at the same time, but it shows you the tokens per second (T/s) for every response, which is the number that matters (and presumably won't include the time spent in the queue).
Here, in fact, the explanation is not adequate. Let's analyze:
> Welcome Groq® Prompster! Are you ready to experience the world's fastest Large Language Model (LLM)?
"The world's fastest LLM? They must have made an LLM, I guess"
> We'd suggest asking about a piece of history, requesting a recipe for the holiday season, or copy and pasting in some text to be translated; "Make it French."
"Instructions, ok".
> This alpha demo lets you experience ultra-low latency performance using the foundational LLM, Llama 2, 70B created by Meta AI - running on Groq.
"So it's not their own LLM? It's just LLaMa? What's interesting about that? And what's Groq, the thing this is 'running on'? The website? I kind of guessed that, by virtue of being here."
> Like any AI demo, accuracy, correctness, or appropriateness cannot be guaranteed.
"Ok, sure".
Nowhere does it say "we've built new, very fast processing hardware for LLMs. Here's a demo a bog-standard LLM, LLaMa 70B, running on our hardware. Notice how fast it is".
Re: Groqchat
#67Earlier quoted context omitted.
Thanks, impressive full-stack work. I'm sure this was named long before Musk decided to set 44B and change on fire but at first I confused it with Twitter's own LLM thing.
They addressed it on their blog https://groq.com/hey-elon-its-time-to-cease-de-grok/
Re: Groqchat
#68Re: Groqchat
#69Re: Groqchat
#70This isn't running on one chip. It's running on 128, or two racks worth of their kit. https://news.ycombinator.com/item?id=38739106 This doesn't mean much without comparing $ or watts of GPU equivalents