Live data from Hacker News

Groq CEO: 'We No Longer Sell Hardware'

eetimes.com

81–90 of 152 posts

Re: Groq CEO: 'We No Longer Sell Hardware'

#81
post #52

Earlier quoted context omitted.

> latency optimized solutions would be faster - in particular for time-to-first-token sensitive applications Do you have any idea how fast Groq is? Go try it. Consistently over 400 t/s for most of the models that they support, and extremely low latency.

time to first token != tokens per second remember that EU -> US is ~150ms unavoidable latency, for example. then your comparison is local H100 vs Grok + 150ms latency to first token.

I'm in Australia. I have 249ms of unavoidable latency and I'd still use the groq API if I could. It's that much faster than other inference solutions.

Re: Groq CEO: 'We No Longer Sell Hardware'

#82
post #79

Interesting, I guess that is why I never got a response back from them about buying their stuff. My guess is that they realized that just selling hardware is a lot harder than running it themselves. Deploying this level of compute is non-trivial, with very high rates of failure, as well as huge supply chain issues. If you have to sell the hardware and support people buying it, that is a world of trouble. > no-one wan…

How's the overall software support for MI300 series? The hardware itself looks great. (also, +100 to valuing honesty and transparency)

The hardware is actually pretty amazing. 192GB (or 1.5TB in a chassis), is a game changer.

I'll let you know once I get my hands on them again. There really isn't enough public information about them at all. So far, my friends at ElioVP [0] have published a blog post. Still with not enough detail for my taste, but I'm pretty sure he is limited by what he can talk about. Luckily, I am not.

I mention in another comment below that my current goal is to get a bunch of people to perform testing on them and then publish blog posts along with open source code. This way, we can start a repository of CI/CD tests to see how things improve with time. ROCm 6.1 is rumored to be quite an improvement.

[0] https://www.evp.cloud/post/diving-deeper-insights-from-our-l...

Re: Groq CEO: 'We No Longer Sell Hardware'

#83
post #79

Earlier quoted context omitted.

How's the overall software support for MI300 series? The hardware itself looks great. (also, +100 to valuing honesty and transparency)

The hardware is actually pretty amazing. 192GB (or 1.5TB in a chassis), is a game changer. I'll let you know once I get my hands on them again. There really isn't enough public information about them at all. So far, my friends at ElioVP [0] have published a blog post. Still with not enough detail for my taste, but I'm pretty sure he is limited by what he can talk about. Luckily, I am not. I mention in another comment…

Nice! I'm pretty interested in GPGPU applications and MI300A, but I'm also just glad for more competition. Love that you hit up the LocalLLaMa sub.

Do you know if anyone's tested CuPy stuff on MI300X?

Re: Groq CEO: 'We No Longer Sell Hardware'

#84

Earlier quoted context omitted.

> If you have to sell the hardware and support people buying it, that is a world of trouble. What is the difference between this and having to sell the cloud access and supporting the people who buy a subscription?

A lot. It's the same reason amazon doesn't sell servers and instead gives you access to a single instance that everyone pretends is the same but in reality is massively transient.

We give full bare metal access. If you just want one GPU, we give you a VM with PCIe pass through. If you take a whole box, we can give BMC and can give you access to the PDU itself, to hard reboot things remotely. It is as if you own the whole machine yourself.

Re: Groq CEO: 'We No Longer Sell Hardware'

#85
post #83

Earlier quoted context omitted.

The hardware is actually pretty amazing. 192GB (or 1.5TB in a chassis), is a game changer. I'll let you know once I get my hands on them again. There really isn't enough public information about them at all. So far, my friends at ElioVP [0] have published a blog post. Still with not enough detail for my taste, but I'm pretty sure he is limited by what he can talk about. Luckily, I am not. I mention in another comment…

Nice! I'm pretty interested in GPGPU applications and MI300A, but I'm also just glad for more competition. Love that you hit up the LocalLLaMa sub. Do you know if anyone's tested CuPy stuff on MI300X?

We haven't spec'd to buy A's quite yet as you're actually the first person I've heard even suggest them. If you're truly interested, hit me up personally.

By default, we are putting dual 9754's in the chassis, along with 3TB ram and 155TB nvme. A pretty beefy box. However, if you want to work with us, we can customize this to whatever customers need.

Effectively, we are the capex/opex for something that requires a lot of upfront funding and want to work with businesses that would rather focus on the software side of things.

Re: Groq CEO: 'We No Longer Sell Hardware'

#86
post #53

Earlier quoted context omitted.

What open source model are you using when you hit groq? I just benchmarked some perf for some of my larger context window queries last week and groq's API took 1.6 seconds versus 1.8 to 2.2 for OpenAI GPT-3.5-turbo. So, it wasn't much faster. I almost emailed their support to see if I was doing something wrong. Would love to hear any details about your workload or the complexity of your queries.

> 1.6 vs 1.8-2.2 seconds I believe certain companies would kill for 20% performance improvements on their main product.

I have lots of questions about how important latency is since you may be replacing many minutes or hours of a person’s time with undoubtedly a quicker response by any measure. This seems like a knee jerk reaction assuming latency is as important as it’s been with advertising.

I’m not convinced latency matters as much as groqs material tries to claim it does.

Re: Groq CEO: 'We No Longer Sell Hardware'

#87
post #35

Custom state-of-the-art silicon is ridiculously expensive. For a minimum 100 wafers = 10k chips, Groq may have paid $100M = $10k/chip purely in amortizing design costs. Chip design (software + engineer time) and fabrication setup (lithography masks) grow exponentially [1][2] with smaller nodes, e.g., maybe $100M for Groq's current 14nm chips to ~$500M for their planned 4nm tapeout. Once you reach mass production (>>1…

> Custom state-of-the-art silicon is ridiculously expensive.

Think about the amount of money being dumped into "AI" at this point. If you've got the technology and people to make stuff faster/better/cheaper, finding investors to dump money into your chip making business is probably not as hard as it was 2 years ago.

Groq is making this change for other reasons than the expense of tapping out chips.

Re: Groq CEO: 'We No Longer Sell Hardware'

#88
post #74

Earlier quoted context omitted.

For what it’s worth, to me, your approach is the one I’d prefer as a customer.

Thank you. I've been on HN since 2009. The most successful people I've seen here, are the ones that are transparent, honest and ethical. I'm not trying to point fingers, I'm just focused on building a sustainable business and listening to my customers needs. The only way I can do that is by communicating with everyone around me as clearly and openly as I can. All our customers will know exactly where they stand, at a…

Just wanted to say that I like how you see things. Also, if you need help with infrastructure automation and are interested in using Nix for increased reproducibility, get in touch. I know you said you want to keep lean and already have a pipeline of potential candidates, but just in case.

Re: Groq CEO: 'We No Longer Sell Hardware'

#89
post #67

Earlier quoted context omitted.

I'll look into it, though seeing "contact us" always makes me think they're not going to sell a single unit to a home user. (With that said, Groq probably wouldn't either. You can technically buy LPUs for 20k each, without an expectation of support, but it takes tens of them to run Mixtral.) Tenstorrent also looks incredibly Python-specific (as in, everything including their SMI seems mostly Python-based) which doesn…

fwiw as a consumer I have a tenstorrent card in my machine

that is great. please hit me up privately, i'd love to chat with you about it.

Re: Groq CEO: 'We No Longer Sell Hardware'

#90
post #53

Earlier quoted context omitted.

What open source model are you using when you hit groq? I just benchmarked some perf for some of my larger context window queries last week and groq's API took 1.6 seconds versus 1.8 to 2.2 for OpenAI GPT-3.5-turbo. So, it wasn't much faster. I almost emailed their support to see if I was doing something wrong. Would love to hear any details about your workload or the complexity of your queries.

> 1.6 vs 1.8-2.2 seconds I believe certain companies would kill for 20% performance improvements on their main product.

Model quality matters a ton too. They aren't serving OpenAI or Anthropic models, which are state of the art.
Post reply on HN