Groq surpasses 1,200 tokens/sec with Llama 3 8B
1–10 of 33 posts
Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#2Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#3Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#4Is groq related to Twitter's grok or is that just a very unfortunate naming coincidence?
Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#5Is groq related to Twitter's grok or is that just a very unfortunate naming coincidence?
Groq - mainly hardware, the LPU (https://wow.groq.com/lpu-inference-engine/)
Grok - Elon's de jour AI endeavor
Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#6My bullshit detector went off when I first saw Groq posted on HN - a startup is making their own chips (doubt) that performs faster than anything Nvidia has for inference (doubt) and accelerates LLMs to hundreds/thousands of tokens per second?? Mega doubt.
But... then I tried their demo, and... yeah, it's that good. Such an amazing company of talented individuals.
Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#7Is groq related to Twitter's grok or is that just a very unfortunate naming coincidence?
Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#8 When will Groq support a real API (not experimental beta preview)?
When will Groq support logprobs?!
When will Groq actually tell us what their rate limit is?!
Until these aren't answered, many of us can't actually build on Groq.Edit: It seems I'm getting downvoted by Groq employees...
Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#9Is groq related to Twitter's grok or is that just a very unfortunate naming coincidence?
Re: Groq surpasses 1,200 tokens/sec with Llama 3 8B
#10When reading Hacker News you develop a signal/noise filter, where lots of headlines make bold claims but you filter them out as embellishment or exaggeration. My bullshit detector went off when I first saw Groq posted on HN - a startup is making their own chips (doubt) that performs faster than anything Nvidia has for inference (doubt) and accelerates LLMs to hundreds/thousands of tokens per second?? Mega doubt. But.…
The other issue they don't mention is power, space, efficiency etc. We want to run larger models with less power, fewer server blades, at lower cost. Not use more server blades, more chips, more power, etc.