Live data from Hacker News

Groqchat

chat.groq.com

81–90 of 131 posts

Re: Groqchat

#81

I can't wait until LLMs are fast enough that a single response can actually be a whole tree of thought/review process before giving you an answer, yet is still fast enough to not even notice

Why wait? This is pretty much what Groq has in hardware, just need the software layer to do the review process.

Re: Groqchat

#82
post #67
post #43

Earlier quoted context omitted.

They addressed it on their blog https://groq.com/hey-elon-its-time-to-cease-de-grok/

can you trademark a verb ?

Seems like they have it registered. I'm sure they've already lawyered up and will protect their trademark. I think they have a pretty strong case. Maybe elon will license it or buy them.

Re: Groqchat

#83
post #26

Lots of comments talking about the model itself. This is Llama 2 70B, a model that has been around for a while now, so we're not seeing anything in terms of model quality (or model flaws) we haven't seen before. What's interesting about this demo is the speed at which it is running, which demonstrates the "Groq LPU™ Inference Engine". That's explained here: https://groq.com/lpu-inference-engine/ > This is the world’s…

this is running on custom hardware, if you’re curious about the underlying architecture check the publication below. https://groq.com/wp-content/uploads/2023/05/GroqISCAPaper202... EDIT: i work at Groq, but i’m commenting in a personal capacity. happy to answer clarifying questions or forward them along to folks who can :)

Will you be selling individual cards? Are you looking for use cases in the healthcare vertical (noticed its not on your current list)? Working in the medical imaging space and could use this tech as part of the offering. Reach out at 16bit.ai

Re: Groqchat

#84

In case, it's not blinding obvious to people. Groq are a hardware company that have built chips that are designed around the training and serving of machine models particularly targeted at LLMs. So the quality of the response isn't really what we're looking for here. We're looking for speed i.e. tokens per second. I actually have a final round interview with a subsidiary of Groq coming up and I'm very undecided as to…

> the quality of the response isn't really what we're looking for here. We're looking for speed i.e. tokens per second.

But if it was generating high-quality responses, would that not make it go slower?

Re: Groqchat

#85

Earlier quoted context omitted.

fast but wrong/gibberish

Its using vanilla llama-2 from Meta with no fine tuning. The point here is the speed and responsiveness of the underlying HW and SW.

But if the quality of the response is poor, it's irrelevant that it was generated quickly. If it was using different data to generate higher quality responses, would that not slow it down?

Re: Groqchat

#86

Yeah, it’s fast but almost always wrong. I asked it a few things (recipes, trivia etc…) and it completely made up the answers. These things don’t really know how to say “I don’t know” and pretend to know everything.

Can you provide specifics for what you asked what it answered? It seems to answer my questions, including recipes correctly.

I asked it to explain several plot points in the TV series "Foundation", and it got them wrong and admitted it when pressed. Several times. Specifically, why does Raych Foss kill Hari Seldon, and some follow-up questions.

Re: Groqchat

#87
For someone who is totally clueless, I can see it's faster than chat gpt in responding to the same question.

What are some relevant speed metrics? Output tokens per second? How about number of input tokens -- does that matter/how does that factor in.

Re: Groqchat

#88

Another censored and boring Google reader. It lied to me twice in 4 prompts and was forced to apologise when called out. Am I wrong in thinking that the first company to develop an unfiltered and genuine intelligence is going to win this AI game?

Yes, you are.

Re: Groqchat

#89

The point isnt that they are running Llama2-70B. The point is that they are running Llama2-70B faster than anyone else so far.

Out of sheer curiosity, why did you make an account for this thread?

at some point each one of us made an account because of a thread

Re: Groqchat

#90

In case, it's not blinding obvious to people. Groq are a hardware company that have built chips that are designed around the training and serving of machine models particularly targeted at LLMs. So the quality of the response isn't really what we're looking for here. We're looking for speed i.e. tokens per second. I actually have a final round interview with a subsidiary of Groq coming up and I'm very undecided as to…

They are putting the whole LLM into SRAM across multiple computing chips, IIRC. That is a very expensive way to go about serving a model, but should give pretty great speed at low batch size.
Post reply on HN