Earlier quoted context omitted.
How do you know that it didn't somehow find the largest prime? Perhaps you just threw away a Noble Prize.
Nobel Prize in what? There is no Nobel in mathematics or computer science.
Groq runs Mixtral 8x7B-32k with 500 T/s
221–230 of 482 posts
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#222Earlier quoted context omitted.
Filled the form for API Access last night. Is there a delay with increased demand now?
Yes, there's a huge amount of demand because Twitter discovered us yesterday. There will be a backlog, so sorry about that.
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#223Sorry, I'm a bit naïve about all of this. Why is this impressive? Can this result not be achieved by throwing more compute at the problem to speed up responses? Isn't the fact that there is a queue when under load just indicative that there's a trade-off between "# of request to process per unit of time" and "amount of compute to put into a response to respond quicker"? https://raw.githubusercontent.com/NVIDIA/Tensor…
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#224Earlier quoted context omitted.
The weights are quantized to FP8 when they're stored in memory, but all the activations are computed at full FP16 precision.
Can you explain if this affects quality relative to fp16? And is mixtral quantized?
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#225Earlier quoted context omitted.
Appreciate the quick reply! That's interesting.
You're welcome. Thanks for reporting. It's pretty confusing so maybe we should change it :)
They allow you to configure chat participants (a model + params like context or temp) and then each AI answers each question independently in-line so you can compare and remix outputs.
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#226Sorry, I'm a bit naïve about all of this. Why is this impressive? Can this result not be achieved by throwing more compute at the problem to speed up responses? Isn't the fact that there is a queue when under load just indicative that there's a trade-off between "# of request to process per unit of time" and "amount of compute to put into a response to respond quicker"? https://raw.githubusercontent.com/NVIDIA/Tensor…
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#227Sorry, I'm a bit naïve about all of this. Why is this impressive? Can this result not be achieved by throwing more compute at the problem to speed up responses? Isn't the fact that there is a queue when under load just indicative that there's a trade-off between "# of request to process per unit of time" and "amount of compute to put into a response to respond quicker"? https://raw.githubusercontent.com/NVIDIA/Tensor…
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#228If the page can't access certain fonts, it will fail to work, while it keeps retrying requests: https://fonts.gstatic.com/s/notosansarabic/[...] https://fonts.gstatic.com/s/notosanshebrew/[...] https://fonts.gstatic.com/s/notosanssc/[...] (I noticed this because my browser blocks these de facto trackers by default.)
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#229Earlier quoted context omitted.
We connect hundreds of chips across several racks with fast interconnect.
How fast is the memory bandwidth of that fast interconnect?
https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...
Re: Groq runs Mixtral 8x7B-32k with 500 T/s
#230Earlier quoted context omitted.
We are prioritising building out whole systems at the moment I don't think we'll have a consumer level offering in the near future.
I will mention: A lot of innovation in this space comes bottom-up. The sooner you can get something in the hands of individuals and smaller institutions, the better your market position will be. I'm coding to NVidia right now. That builds them a moat. The instant I can get other hardware working, the less of a moat they will have. The more open it is, the more likely I am to adopt it.