Live data from Hacker News

100x defect tolerance: How we solved the yield problem

cerebras.ai

171–180 of 186 posts

Re: 100x defect tolerance: How we solved the yield problem

#172
post #2

I think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't nee…

Of course many people are going to collectively lose trillions, AI's a very highly hyped industry with people racing into it without an intellectual edge and any temporary achievement by any one company will be quickly replicated and undercut by another using the same tools. Economic success of the individuals swarming on a new technology is not a guarantee whatsoever, nor is it an indicator of the impact of the tech…

Dollars are not lost; they are just very indirectly invested into gpu makers (and energy providers)

Re: 100x defect tolerance: How we solved the yield problem

#173
post #159
post #62

Earlier quoted context omitted.

Redundant cores lead to a fault tolerant chip.

ECC memory is fault tolerant. It repairs issues on the fly without disabling hardware. This on the other hand is merely redundant to handle manufacturing defects. If they make a mistake and ship a bad core that malfunctions at runtime, it is not going to tolerate that.

Redundancy is a method of providing fault tolerance, the existence of other methods doesn't make it less fault tolerant.

Nothing is tolerant to all possible faults. Fault tolerance refers to being able to tolerate specific types of faults under specific conditions.

Fault tolerant is the proper term for this.

Re: 100x defect tolerance: How we solved the yield problem

#174
post #145

Earlier quoted context omitted.

I did eventually get an LLM to produce what seems to be a correct diagram of a sentence it had never seen, but it took about ten tries. Grammatical analysis seems to have happened correctly every time, but getting to a usable diagram was difficult. (I know that it's generally rude to include LLM output in HN comments, but in this case I think it's essential supporting material to elevate the discussion of LLM capabil…

(I know that it's generally rude to include LLM output in HN comments, but in this case I think it's essential supporting material to elevate the discussion of LLM capabilities above "yes it is", "no it isn't".) You just have to be prepared to take a karma hit for it. The audience here does not consist largely of 'hackers', but seems to skew toward the sort of fearful, resentful reactionaries that hacker culture trad…

ASCII art Reed–Kellogg sentence diagrams are probably hard to find anywhere, and Graphviz can't really express Reed–Kellogg diagrams. But Reed and Kellogg published their somewhat ad-hoc diagram language in 01877, 78 years before what we now call "linguistics" was known in the West thanks to Chomsky's work in 01955. These are among the reasons I thought it might be a good idea to use the form of sentence diagrams used by linguists instead of the more compact Reed–Kellogg diagrams.

Re: 100x defect tolerance: How we solved the yield problem

#175
post #39

I have a dumb question. Why isn't silicon sold in cubes instead of cylinders?

The silicon ingots have a rotating production process that results in cylinders, not bricks.

fascinating, I figured it was something like that. maybe we should produce hexagonal, instead of square, chip designs

Re: 100x defect tolerance: How we solved the yield problem

#176
post #173
post #159

Earlier quoted context omitted.

ECC memory is fault tolerant. It repairs issues on the fly without disabling hardware. This on the other hand is merely redundant to handle manufacturing defects. If they make a mistake and ship a bad core that malfunctions at runtime, it is not going to tolerate that.

Redundancy is a method of providing fault tolerance, the existence of other methods doesn't make it less fault tolerant. Nothing is tolerant to all possible faults. Fault tolerance refers to being able to tolerate specific types of faults under specific conditions. Fault tolerant is the proper term for this.

I think it would have been better to write redundant. It is more specific.

Re: 100x defect tolerance: How we solved the yield problem

#177
post #163
post #160

Earlier quoted context omitted.

I’m not so sure it’s going to even do that much. People are currently happy to use LLM’s, but the outputs aren’t accurate and don’t seem to be improving quickly. A YouTuber watch regularly includes questions they asked Chat GPT and very single time there’s a detailed response in the comments showing how the output is wildly wrong from multiple mistakes. I suspect the backlash from disgruntled users is going to hit th…

Using function calls for correct answer lookup already practically eliminates this, it's not wide spread yet, but the ease of doing it is already practical for many. New models aren't being trained specifically on single answers which will only help. The expense for the larger models is something to be concerned about. Small models with function calls is already great, especially if you narrow down what they are bein…

Right. And any particular question people think AIs are bad at also has a comments section of people who have run better crafted prompts that do the job just fine. The consensus is heading more towards "well damn, actually LLMs might be all we need" rather than "LLMs are just a stepping stone" - but either way, that's fine, cuz plenty of more advanced architecture uses are on their way (especially error correction / consistency frameworks).

I dont believe there are any significant academic critiques doubting this. There are a lot of armchair hot takes, and perceptions that this stuff isn't improving up to their expectations, but those are pretty divorced from any rigorous analysis of the field, which is still improving at staggeringly fast rates compared to any other field of research. Aint no wall, folks.

Re: 100x defect tolerance: How we solved the yield problem

#178
post #163

Earlier quoted context omitted.

Using function calls for correct answer lookup already practically eliminates this, it's not wide spread yet, but the ease of doing it is already practical for many. New models aren't being trained specifically on single answers which will only help. The expense for the larger models is something to be concerned about. Small models with function calls is already great, especially if you narrow down what they are bein…

Right. And any particular question people think AIs are bad at also has a comments section of people who have run better crafted prompts that do the job just fine. The consensus is heading more towards "well damn, actually LLMs might be all we need" rather than "LLMs are just a stepping stone" - but either way, that's fine, cuz plenty of more advanced architecture uses are on their way (especially error correction /…

“Crafting a better prompt” is often simply spinning an RNG again and again until you end up with an answer that happens to be good enough.

In the real world if you know the correct answer you don’t need to ask the question. A self driving car that needs you to pay attention isn’t self driving.

Any system can get canned response, the value of AI is completely in its ability to handle novelty without hand holding. And none of these systems actually do that even vaguely well in practice rather than providing response that are vaguely close to correct.

If I ask for a summary of an article and it gets anything wrong in the article that’s a 0 because now I need to read the article to know what it said. Arguably the value is actually negative here.

Re: 100x defect tolerance: How we solved the yield problem

#179
Ever heard the old joke story about an American buyer told the Japanese manufacture how many incorrectly made bolts were acceptable per lot of a thousand bolts? Maybe 2 or 3 in 1,000?

So the Japanese didn't have any incorrectly made bolts in their manufacturing process so they just added two or three bad ones to every batch to please the Americans.

Re: 100x defect tolerance: How we solved the yield problem

#180

Earlier quoted context omitted.

> And just as every other hype cycle, this one will crash down hard. Isn't that an inherent problem with pretty much everything nowadays: crypto, blockchain, AI, even the likes of serverless and Kubernetes, or cloud and microservices in general. There's always some hype cycle where the people who are early benefit and a lot of people chasing the hype later lose when the reality of the actual limitations and the real…

> I don't think the current "AI" is special in any way As someone who loves to pour ice water on AI hype, I have to say: you can't be serious. The current AI tech has opened up paths to develop applications that were impossible just a few years ago. Even if the tech freezes in place, I think it will yield substantial economic value in the coming years. It's very different from crypto, the main use case for which appe…

> Even if the tech freezes in place, I think it will yield substantial economic value in the coming years.

The question is, where will this "economic value" be? Because "economic value" and actual progress that helps society are two very different things. For example, if someone wants to hire people, they can use "AI" to sift through the applications. But people looking for a job can also use "AI" to write their applications. In the end you may have created "economic value", but its an arms race and a waste of resources at the core, more digital paperwork, a waste of compute. So the actual value is not positive, it is negative. And we see that in many places where this so called "AI" is supposed to help.

Does it mean it is entirely useless? No, but the field of applications where it actually makes sense and has an overall net benefit is way smaller than many believe.

Plus, there are different types of neural networks in use for decades already. Look at OCR, for example, where the commercial OCR software switched to neural networks around the mid 90s already. So it is not that "AI" as such is bad, just that this generative neural network stuff is overly hyped by people who have absolutely no clue about it, but have to hop on the bandwagon to not be left out and keep shareholder values up, because most of these shareholders are equally stupid about the whole issue. Its a circus that burns many resources, money that could have created way more value in other areas.

Post reply on HN