Live data from Hacker News

100x defect tolerance: How we solved the yield problem

cerebras.ai

81–90 of 186 posts

Re: 100x defect tolerance: How we solved the yield problem

#81
post #75
post #2

I think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't nee…

"While I continue to believe that many people are going to collectively lose trillions of dollars ultimately pursuing "AI" at this stage" Can you please explain more why you think so ? Thank you.

It's a hype cycle with many of the hypers and deciders having zero idea about what AI actually is and how it works. ChatGPT, while amazing, is at its core a token predictor, it cannot ever get to an AGI level that you'd assume to be competitive to a human, even most animals.

And just as every other hype cycle, this one will crash down hard. The crypto crashes were bad enough but at least gamers got some very cheap GPUs out of all the failed crypto farms back then, but this time so much more money, particularly institutional money, is flowing around AI that we're looking at a repeat of Lehman's once people wake up and realize they've been scammed.

Re: 100x defect tolerance: How we solved the yield problem

#82

Earlier quoted context omitted.

If rack mounted, you are ending up with something like a reverse power station. So why not use it as an energy source? Spin a turbine.

I'm aware of the efficiency losses but I think it would be amusing to use that turbine to help power the machine generating the heat.

Hey, we're building artificial general intelligence, what's a little perpetual motion on the side?

Re: 100x defect tolerance: How we solved the yield problem

#83

Earlier quoted context omitted.

If you let the chip actual boil enough water to run a turbine you're going to have a hard time keeping the magic smoke inside. Much better to run at reasonable temps and try to recover energy from the waste heat.

What if you chose a refrigerant with a lower boiling point?

That's basically the principle of binary cycle[1] generators. However for data center waste heat recovery, I'd think you'd want to use a more stable fluid for cooling, and then pump it to a separate closed-loop binary-cycle generator. No reason to make your datacenter cooling system also deal with high pressure fluids, and moving high pressure working fluid from 1000s of chips to a turbine of sufficient size, etc.

[1]: https://en.wikipedia.org/wiki/Binary_cycle

Re: 100x defect tolerance: How we solved the yield problem

#84
post #2

I think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't nee…

> Xilinx was still aggressively suing people who put SERDES ports on FPGAs

This so isn't important to your overall point, but where would I begin to look into this? Sounds fascinating!

Re: 100x defect tolerance: How we solved the yield problem

#85

This is a strange blog post. Their tables say: Cerebras yields 46225 * .93 = 43000 square millimeters per wafer NVIDIA yields 58608 * .92 = 54000 square millimeters per wafer I don't know if their numbers are correct but it is a strange thing for a startup to brag that it is worse than a big company at something important.

Being within striking distance of SOTA while using orders of magnitude fewer resources is worth bragging about.

Re: 100x defect tolerance: How we solved the yield problem

#86

Earlier quoted context omitted.

Any thoughts on why they are disabling so many cores in their current product? I did some quick noodling based on the 46/970000 number and the only way I ended up close to 900,000 was by assuming that an entire row or column would be disabled if any core within it was faulty. But doing that gave me a ~6% yield as most trials had active core counts in the high 800,000s

They did mention that they stash extra cores to enable the re-routing. Those extra cores are presumably unused when not routed in.

That was my first thought but based on the rerouting graphic it seems like the extra cores would be one or two rows and columns around the border which would only account for ~4000 cores.

Re: 100x defect tolerance: How we solved the yield problem

#87

Neat. What about power density? An H100 has a TDP of 700 watts (for the SXM5 version). With a die size of 814 mm^2 that's 0.86 W/mm^2. If the cerebras chip has the same power density, that means a cerebras TDP of 37.8 kW. That's a lot. Let's say you cover the whole die area of the chip with water 1 cm deep. How long would it take to boil the water starting from room temperature (20 degrees C)? amount of water = (die…

The enthalpy of vaporization of water (at standard pressure) is listed by Wikipedia[1] as 2.257 kJ/g, so boiling 462 grams would require an additional 1.04 MJ, adding 26 seconds. Cerebras claims a "peak sustained system power of 23kW" for the CS-3 16 Rack Unit system[2], so clearly the power density is lower than for an H100. [1] https://en.wikipedia.org/wiki/Enthalpy_of_vaporization#Other... [2] https://cerebras.ai/…

On a tangent: has anyone built an active cooling system which operates in a partial vacuum? At half atmospheric pressure, water boils at around 80 C, which i believe is roughly the operating temperature for a hard-working chip. You could pump water onto the chip, have it vapourise, taking away all that heat, then take the vapour away and condense it at the fan end.

This is how heat pipes work, i believe, but heat pipes aren't pumped, they rely entirely on heat-driven flow. I would have thought there were pumped heat pipes. Are they called something else?

It's also not a refrigerator, because those use a pump to pressurise the coolant in its gas phase, whereas here you would only be pumping the water.

Re: 100x defect tolerance: How we solved the yield problem

#88
post #4

So they massively reduce the area lost to defects per wafer, from 361 to 2.2 square mm. But from the figures in this blog, this is massively outweighed by the fact that they only get 46222 sq mm useable area out of the wafer, as opposed to 56247 that the H100 gets - because they are using a single square die instead of filling the circular wafer with smaller square dies, they lose 10,025 sq mm! Not sure how that's a…

Why does their chip have to be rectangular, anyways? Couldn't they cut out a (blocky) circle too?

[deleted]

Re: 100x defect tolerance: How we solved the yield problem

#89
post #75
post #2

I think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't nee…

"While I continue to believe that many people are going to collectively lose trillions of dollars ultimately pursuing "AI" at this stage" Can you please explain more why you think so ? Thank you.

I would guess you're not asking a serious question here but if you were feel free to contact me, it's why I put my email address in my profile.

Re: 100x defect tolerance: How we solved the yield problem

#90
post #2

I think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't nee…

Any thoughts on why they are disabling so many cores in their current product? I did some quick noodling based on the 46/970000 number and the only way I ended up close to 900,000 was by assuming that an entire row or column would be disabled if any core within it was faulty. But doing that gave me a ~6% yield as most trials had active core counts in the high 800,000s

I could guess that it helps with heat dissipation/management. But I don't know. That guess is from looking at the list of patents[1] they have.

[1] https://patents.justia.com/assignee/cerebras-systems-inc

Post reply on HN