Live data from Hacker News

How to scale your model: A systems view of LLMs on TPUs

jax-ml.github.io

1–10 of 31 posts

Re: How to scale your model: A systems view of LLMs on TPUs

#5
post #4

I am really looking forward for JAX to take over pytorch/cuda over the next years. The whole PTX kerfuffle with Deepseek team shows the value of investing in more low levels approaches to squeeze out the most out of your hardware.

Most Pytorch users don’t bother even with the simplest performance optimizations, and you are talking about PTX.

Re: How to scale your model: A systems view of LLMs on TPUs

#10

An author's tweet thread: https://x.com/jacobaustin132/status/1886844716446007300

Here in the thread he says: https://x.com/jacobaustin132/status/1886844724339675340 : `5 years ago, there were many ML architectures, but today, there is (mostly) only one [transformers].`

To what degree is this actually true, and what else is on the horizon that might become as popular as transformers?

Post reply on HN