But what are these LPUs optimized for: tensor operations (like Google's TPUs) or LLMs/Transformers architecture?
If it is the latter, how would they/their clients adapt if a new (improved) architecture hits the market?
31–40 of 131 posts
But what are these LPUs optimized for: tensor operations (like Google's TPUs) or LLMs/Transformers architecture?
If it is the latter, how would they/their clients adapt if a new (improved) architecture hits the market?
Lots of comments talking about the model itself. This is Llama 2 70B, a model that has been around for a while now, so we're not seeing anything in terms of model quality (or model flaws) we haven't seen before. What's interesting about this demo is the speed at which it is running, which demonstrates the "Groq LPU™ Inference Engine". That's explained here: https://groq.com/lpu-inference-engine/ > This is the world’s…
https://groq.com/wp-content/uploads/2023/05/GroqISCAPaper202...
EDIT: i work at Groq, but i’m commenting in a personal capacity.
happy to answer clarifying questions or forward them along to folks who can :)
This doesn't mean much without comparing $ or watts of GPU equivalents
This isn't running on one chip. It's running on 128, or two racks worth of their kit. https://news.ycombinator.com/item?id=38739106 This doesn't mean much without comparing $ or watts of GPU equivalents
More info about Groq: https://groq.com/lpu-inference-engine/
The point isnt that they are running Llama2-70B. The point is that they are running Llama2-70B faster than anyone else so far.
[flagged]