Has anyone been running LLMs on TPUs in prod? Curious to hear experiences.
Haven't been too impressed with inference versus tensor rt llm for example though
51–60 of 66 posts
Has anyone been running LLMs on TPUs in prod? Curious to hear experiences.
Haven't been too impressed with inference versus tensor rt llm for example though
This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…
Genesis Cloud started integration and testing of Gaudi2 quite a while ago. I fully agree with the take of the article.
I can't promise per hour rental, but for longer times they are available! (should you be interested you can find contact details on the website)
One question I have that nobody, including an Intel AXG employee, has been able to answer satisfsctorily for me is why both Gaudi and Ponte Vecchio exist. Wouldn't Intel have better chances of success if they focused on one product line?
This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…
Disclaimer: I technically still am employed at Genesis Cloud (though no longer actively involved). Genesis Cloud started integration and testing of Gaudi2 quite a while ago. I fully agree with the take of the article. I can't promise per hour rental, but for longer times they are available! (should you be interested you can find contact details on the website)
This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…
Earlier quoted context omitted.
I can feel the NVDA stock slipping as we speak… It has been amazing watching the groupthink at work on that stock when we just saw the same group do it on TSLA to disastrous effect. A similar no moat situation where they simply can’t imagine competitors ever existing.
The Model Y is the best selling car in the world in 2023. Those of us who were buying in 2019 are still up quite a bit even though the stock was higher at one point. RIVN, Ford, GM all are losing a lot of money on every EV they sell. We were right to bet on TSLA being a major winner. I actually put 40% of my TSLA into NVDA last year, because the demand for AI hardware is going to keep going up. I'm not saying the sto…
NVIDIA's profit margin is almost 92% on an H100. I'm surprised more chip companies haven't jumped on a "ML accelerator" bandwagon by now.
There's a dozen AI chips already; how many do you want? Now working ones is a different story.
Just curious because IME that's the point where the fun problems surface :)
Earlier quoted context omitted.
There's a dozen AI chips already; how many do you want? Now working ones is a different story.
For anyone deeply familiar with building one: what are the biggest problems you run into 3 months in that you didn't foresee? Just curious because IME that's the point where the fun problems surface :)
This is nice to foster some competition in hardware for model training, but the availability of these machines seems very limited - I don't think there's any major cloud provider allowing per hour rental of Gaudi2 VMs and Intel's own site directs you to buy an 8x GPU provisioned server from Supermicro for more than 40k USD. Availability and software stack is still heavily in Nvidia's favor right now, but maybe by the…
Think it'll probably crack on with Gaudi3 at 4x performance, twice VRAM etc later this year. We found cuda sycl conversion surprisingly good https://www.intel.com/content/www/us/en/developer/articles/t...