Live data from Hacker News

An inside look at the custom CPUs in Tesla's Dojo Supercomputer

semianalysis.com

31–40 of 132 posts

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#31
> If their claims are true, Tesla has 1 upped everyone in the AI hardware and software field. I’m skeptical, but this is also a hardware geek’s wet dream.

How is the author declaring Tesla has one upped everyone in AI hardware and software while the article has exactly zero references to TPUs?

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#34
post #31

> If their claims are true, Tesla has 1 upped everyone in the AI hardware and software field. I’m skeptical, but this is also a hardware geek’s wet dream. How is the author declaring Tesla has one upped everyone in AI hardware and software while the article has exactly zero references to TPUs?

Also, these chips are as yet only in Tesla’s labs, not in production. What’s cooking in secret in NVIDIA’s labs right now?NVIDIA doesn’t have to hype their stuff, so they don’t tip their hand to competitors until it’s ready for sale.

Creating a lab beast is one thing, making it useful and economical in production is another.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#35

Insane, truly insane. Meanwhile laptop CPUs, with the notable exception of M1, are impossible to cool.

Not sure that's relevant here - the reason they can cool this is because they're custom-designing housing, server racks, etc for these chips based on the amount of power they draw. You could cool pretty much anything if you were able to give it this much love and care.

Also, the reason most laptops run hot is that most modern high performance processors are thermally limited. It means the cooling is never "sufficient" because if you make the cooling better then you get a faster processor instead of a cooler laptop.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#36
post #31

> If their claims are true, Tesla has 1 upped everyone in the AI hardware and software field. I’m skeptical, but this is also a hardware geek’s wet dream. How is the author declaring Tesla has one upped everyone in AI hardware and software while the article has exactly zero references to TPUs?

It's also a fictional product at this point. Even if the hardware existed, if you can't program it, it won't succeed.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#37
Taking everything at face value: Dojo is overall a very impressive project!

- Communication speed is one of the biggest bottlenecks to large models, so their bandwidth of 4TBps is very smart.

- They claim a 1.3x perf/watt improvement, which is not really that great for ASICs compared to GPUs. Perf/watt is probably the most important number in datacenters.

- They only use SRAM, no DRAM. This is a huge mistake, which limits their model size. You can only fit a ~10GB model inside a single tile, versus 80GB models for a single A100 GPU.

- Software / compiler stack is as or more important than the hardware itself, because it dictates how much real performance you can squeeze out of the chips. I think Tesla will need to heavily focus on this area before getting anywhere close to real-world GPU performance.

Overall, I imagine the project will have similar pitfalls to Cerebras.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#38

Earlier quoted context omitted.

Moore’s law was about count, not density.

It was about cost. https://hasler.ece.gatech.edu/Published_papers/Technology_ov...

Yes, count for optimum cost per transistor. So perhaps GP's statement was incomplete but not wrong.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#40
post #37

Taking everything at face value: Dojo is overall a very impressive project! - Communication speed is one of the biggest bottlenecks to large models, so their bandwidth of 4TBps is very smart. - They claim a 1.3x perf/watt improvement, which is not really that great for ASICs compared to GPUs. Perf/watt is probably the most important number in datacenters. - They only use SRAM, no DRAM. This is a huge mistake, which l…

Yeah, a lot of that probably doesn't matter for Tesla's current internal use case but all of it matters when you're talking about commercializing it.
Post reply on HN