Live data from Hacker News

Groq surpasses 1,200 tokens/sec with Llama 3 8B

twitter.com

21–30 of 33 posts

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#21
post #17
post #16

Earlier quoted context omitted.

Classic HN – downvotes without explanation. I might well be wrong about the etymology here, but I understand "grokking" to be a term for a phenomenon in training neural networks. What I'm not sure about is which was there first – AI companies called some version of "grok" or that term.

The term grok came from Robert Heinlein’s 1961 novel Stranger in a Strange Land and got picked up by the CS field heavily around the late 60s. https://en.wikipedia.org/wiki/Grok

I do know that meaning of "grok", but I always assumed the more specific one in the context of neural networks was what informed these two naming choices, although I really don't know the exact timeline.

Didn't know about Heinlein coining it though, that's cool!

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#22
post #16

Earlier quoted context omitted.

Classic HN – downvotes without explanation. I might well be wrong about the etymology here, but I understand "grokking" to be a term for a phenomenon in training neural networks. What I'm not sure about is which was there first – AI companies called some version of "grok" or that term.

From the HN commenting guidelines > Please don't comment about the voting on comments. It never does any good, and it makes boring reading.

I'm well aware of the "complaints about downvotes beget downvotes" meme and was expecting it here, but sometimes I am genuinely curious about the nature of the disagreement. Here I really just wanted to learn what people think the actual etymology is. I get and appreciate "I don't find this contribution helpful", but I really dislike a "I think you're factually wrong but can't be bothered to correct you" downvote.

As an aside, I wonder when "please don't make a quote from the HN commenting guidelines the only contribution of your comment" will join that list...

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#23
post #17
post #16

Earlier quoted context omitted.

Classic HN – downvotes without explanation. I might well be wrong about the etymology here, but I understand "grokking" to be a term for a phenomenon in training neural networks. What I'm not sure about is which was there first – AI companies called some version of "grok" or that term.

The term grok came from Robert Heinlein’s 1961 novel Stranger in a Strange Land and got picked up by the CS field heavily around the late 60s. https://en.wikipedia.org/wiki/Grok

Unrelated, but you just reminded me of an old blog: Groklaw. I can't believe it's been over 10 years since it was active

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#24
post #22

Earlier quoted context omitted.

From the HN commenting guidelines > Please don't comment about the voting on comments. It never does any good, and it makes boring reading.

I'm well aware of the "complaints about downvotes beget downvotes" meme and was expecting it here, but sometimes I am genuinely curious about the nature of the disagreement. Here I really just wanted to learn what people think the actual etymology is. I get and appreciate "I don't find this contribution helpful", but I really dislike a "I think you're factually wrong but can't be bothered to correct you" downvote. As…

> As an aside, I wonder when "please don't make a quote from the HN commenting guidelines the only contribution of your comment" will join that list...

HN is largely community driven moderation, helping dang do his job, so I suspect this meta don't wouldn't make it

I didn't comment on the substance because others already had by then, not sure why they didn't prefer your OG comment...

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#25
post #2

Groq is an insane company. SambaNova (discussed yesterday[0]) is also very promising. However, what I really want to see is local AI accelerator chips a la Tenstorrent Grayskull that can boost local generation to hundreds of tokens per second while being more efficient than GPUs. [0]: https://news.ycombinator.com/item?id=40508797

Samba is on gen 4 silicon and still lagging, somebody over there is doing something wrong

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#26

When reading Hacker News you develop a signal/noise filter, where lots of headlines make bold claims but you filter them out as embellishment or exaggeration. My bullshit detector went off when I first saw Groq posted on HN - a startup is making their own chips (doubt) that performs faster than anything Nvidia has for inference (doubt) and accelerates LLMs to hundreds/thousands of tokens per second?? Mega doubt. But.…

8 year old unicorn++ with a public demo sounds credible?

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#27
post #15

Earlier quoted context omitted.

> though they deserve every right to protect their trademark. And they can, Twitter (why everything gets claimed as his personal work I never know) isn't using their trademark. As they (Groq) themselves have said... > the difference of one consonant (q, k) only matters to scrabblers and spell checkers Grok the term has been around since at least 1961.[0] The fact that a company decided to take a common term (especial…

Trademarks are more nuanced than you are relaying here. Groq, in arguing that their mark is different from "grok" (at the USPTO) is because one cannot trademark common words. They are applying for plain marks (without font/color/logo) and this is very normal. I went through this with a proper name trademark In the Groq vs Grok, they are arguing that the average person will confuse the marks (as can be seen in many HN…

> Their argument is that Grok should not be given a trademark beforehand due to this potential confusion.

Groq says no such thing. Their two public things so far include

1) a company that rebranded to Groq Healthcare 2) a C&D to twitter over the name

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#28
post #2

Groq is an insane company. SambaNova (discussed yesterday[0]) is also very promising. However, what I really want to see is local AI accelerator chips a la Tenstorrent Grayskull that can boost local generation to hundreds of tokens per second while being more efficient than GPUs. [0]: https://news.ycombinator.com/item?id=40508797

Samba is on gen 4 silicon and still lagging, somebody over there is doing something wrong

How are they lagging? They are running faster than anyone else at full precision and with many many fewer chips than Groq. Groq is not real.

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#29

Is groq related to Twitter's grok or is that just a very unfortunate naming coincidence?

I think groq has more users and a better business model.

Groq makes zero revenue, needs hundreds on chips to run 1 model, and runs everything at lower precision. SambaNova has a lot of revenue, and runs at that speed at full precision on a single node. It really isn’t a competition.

Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B

#30

Earlier quoted context omitted.

The issue is that their chips need a huge amount of server blades and there's a big doubt whether this model actually scales. That is, how will Groq handle much larger models with a context of hundreds of thousands or millions of tokens? Right now this would require them to deploy a cluster with thousands of chips, versus 10 chips for say an NVidia system. The other issue they don't mention is power, space, efficienc…

Cerebrus faces similar challenges with their wafer scale chips. If anything, Google's TPU advancements chart a viable course. I suspect both Groq and Cerebrus will overcome the challenges and offer competitive compute options, depending on the context

SambaNova is the only one if the chip startups that is viable. It surprises me that people don’t see this.
Post reply on HN