Live data from Hacker News

100x defect tolerance: How we solved the yield problem

cerebras.ai

141–150 of 186 posts

Re: 100x defect tolerance: How we solved the yield problem

#141
post #2

I think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't nee…

Of course many people are going to collectively lose trillions, AI's a very highly hyped industry with people racing into it without an intellectual edge and any temporary achievement by any one company will be quickly replicated and undercut by another using the same tools. Economic success of the individuals swarming on a new technology is not a guarantee whatsoever, nor is it an indicator of the impact of the technology.

Just like the dotcom bubble, AI is gonna hit, make a few companies stinking rich, and make the vast majority (of both AI-chasing and legacy) companies bankrupt. And it's gonna rewire the way everything else operates too.

Re: 100x defect tolerance: How we solved the yield problem

#143

Earlier quoted context omitted.

They did mention that they stash extra cores to enable the re-routing. Those extra cores are presumably unused when not routed in.

That was my first thought but based on the rerouting graphic it seems like the extra cores would be one or two rows and columns around the border which would only account for ~4000 cores.

If the system were broken down into more subdivisions internally, there would be more cores dedicated to replacement. It seems like it could be more difficult to reroute an entire row or column of cores on a wafer than a small block. Perhaps, also, they are building in heavy redundancy for POC and in the future will optimize the number of cores they expect to lose.

Re: 100x defect tolerance: How we solved the yield problem

#144
post #79
post #40

Understanding that there's inherent bias by them being competitors of the other companies, but still this article seems to make some stretches. If you told me you had an 8% core defect rate reduced 100x, I'd assume you got to close to 99% enablement. The table at the end shows... Otherwise. They also keep flipping between cores, SMs, dies, and maybe other block sizes. At the end of the day I'm not very impressed. The…

I think you're missing the point. The comparison is not between 93% and 92%. The comparison is between what they're getting (93%) and what you'd get if you scaled up the usual process to the core size they're using (0%). They are doing something different (namely: a ~whole wafer chip) that isn't possible without massively boosting the intra-chip redundancy. (The usual process stops working once you no longer have any…

I think I'll still stand by my viewpoint. They said:

> On the Cerebras side, the effective die size is a bit smaller at 46,225mm2. Applying the same defect rate, the WSE-3 would see 46 defects. Each core is 0.05mm2. This means 2.2mm2 in total would be lost to defects.

So ok they claim that they should see (46225-2.2)/46225 = 99.995%. Doing the same math for their Nvidia numbers it's 99.4%. And yet in practice neither approach got to these numbers. Nowhere near it. I just feel like the whole article talks about all this theory and numbers and math of how they're so much better but in practice it's meaningless.

So what I'm not seeing is why it'd be impossible for all the H100s on a wafer to be interconnected and call it a day. You'd presumably get 92/93 = 98.9% of the performance and, here's the kicker, no need to switch to another architecture. I didn't know where your 0% number came from. Nothing about this article says that a competitor doing the same scaling to wafer scale would get 0%, just a marginal decrease in how many cores made it through fab.

Fundamentally I am not convinced from this article that Cerebras has done something in their design that makes this possible. All I'm seeing is that it'd perform 1% faster.

Edit: thinking a bit more on it, to me it's like they said TSMC has a guy with a sledgehammer who smashes all the wafers and their architecture snaps a tiny bit cleaner. But they haven't said anything about firing the guy with the sledgehammer. Their paragraph before the final table says that this whole exercise is pretty much meaningless because their numbers are made up about competitors and they aren't even the right numbers to be using. Then the table backs up my paraphrase.

Re: 100x defect tolerance: How we solved the yield problem

#145

Earlier quoted context omitted.

it cannot ever get to an AGI level that you'd assume to be competitive to a human, even most animals. Suppose you turn out to be wrong. What would convince you?

It could diagram a sentence it had never seen.

I did eventually get an LLM to produce what seems to be a correct diagram of a sentence it had never seen, but it took about ten tries. Grammatical analysis seems to have happened correctly every time, but getting to a usable diagram was difficult.

(I know that it's generally rude to include LLM output in HN comments, but in this case I think it's essential supporting material to elevate the discussion of LLM capabilities above "yes it is", "no it isn't".)

First, I just asked ChatGPT (GPT-4o mini, what they offer anonymous free users):

> Please diagram the sentence, "If my mother's road had been cultural, my uncle's sloth squad residue would have compiled a periodic meadow of night."

I just invented that sentence, picking several words with a random number generator, but of course I can't prove to you that I didn't copy it from some book you haven't seen. (For that, you'd have to invent a sentence yourself, or at least provide me with some words to include in it.) I hope you can at least agree that it's a highly unlikely sentence for anyone ever to have uttered before!

After talking to itself for a few paragraphs, it arrived at the following grammatical analysis:

  Step 4: Diagramming the sentence

    Main Clause:
        Subject: "my uncle's sloth squad residue"
        Verb: "would have compiled"
        Direct Object: "a periodic meadow of night"
            "a periodic meadow" (noun phrase)
            "of night" (prepositional phrase modifying "meadow")

    Conditional Clause:
        Subject: "my mother's road"
        Verb: "had been" (linking verb)
        Predicate Adjective: "cultural"
This is exactly correct. From that it produced the following ASCII art "diagram", which looks like nonsense to me, but maybe it's in some kind of sentence-diagramming formalism I haven't seen:

                If
        ______/______
       |              |
    my mother's    had been
      road             |
                        cultural
       
   my uncle's sloth squad residue
                 |
    would have compiled
                 |
  a periodic meadow of night
              |
            of
            night
I told it the diagram looked like nonsense, and it produced this one, which looks worse:

                 If
           _______/______
          |               |
  my mother's road    had been
           |             |
       cultural      (Main Clause)
                       |
       my uncle's sloth squad residue
                           |
                 would have compiled
                           |
               a periodic meadow
                         |
                       of night
I asked GPT-4 (paid) the same question, and it gave me another exactly correct grammatical analysis in words:

  1. **Identify the main clauses**: 
     - "If my mother's road had been cultural" is a dependent clause (specifically, a conditional clause).
     - "my uncle's sloth squad residue would have compiled a periodic meadow of night" is the main clause.

  2. **Start with the main clause**:
     - Subject: "my uncle's sloth squad residue"
       - "my uncle's" modifies "sloth squad," which in turn modifies "residue."
     - Predicate: "would have compiled"
       - This includes the modal auxiliary verb ("would"), perfect auxiliary ("have"), and main verb ("compiled").
     - Direct object: "a periodic meadow of night"
       - “a” is an article modifying “meadow,”
       - “periodic” modifies “meadow,”
       - “of night” is a prepositional phrase modifying “meadow.”

  3. **Diagram the dependent clause**:
     - Subject: "road"
       - Modified by possessive noun phrase “my mother’s”
     - Predicate: "had been"
     - Complement/Adjective: "cultural"
(I feel like this is definitely "a level that you'd assume to be competitive to a human", in mschuster91's phrase; most native English speakers can't do that.)

But then it came up with the following incorrect diagram:

                            +---------------------------------------------------+
                            |                      Main Clause                   |
                            +---------------------------------------------------+
                                    /                |               \
                                   /                 |                \
                                  /                  |                 \
                         [My uncle's]         [sloth squad]          [residue]
                             \                    /                     /
                              \                  /                     /
                               \                /                     /
                                would have compiled -------------------
                                                  \
                                                   \
                                                    \
                                                     [a periodic meadow of night]
                                                      /           |            \
                                                     /            |             \
                                                    a        periodic       [of night]

  +---------------------+
  |  Dependent Clause   |
  +---------------------+
        If
         |
        road
      ___|_____
    |          |
   my mother's had been cultural
Telling it the diagram was nonsense didn't help.

But my conclusion from this is not that GPT-4 can't analyze the grammatical structure of sentences it had never seen. It's that GPT-4 gets "confused" when it tries to draw ASCII art, like a person trying to tell you how to draw it over the telephone, but not quite as bad. I've seen similarly abominable results when it tried to draw circuit schematics and guitar tablatures as ASCII art. I tried asking it for a graphviz diagram, which was syntactically correct graphviz code but didn't express the sentence structure properly.

I tried simplifying the problem with the prompt, 'Please provide a syntactic tree showing the syntactic derivation of the sentence, "My uncle's sloth squad residue would have compiled a periodic meadow of night," using graphviz syntax and with individual words at the leaf nodes of the tree.' which yielded a slightly incorrect diagram. (I'm still sure nobody has ever uttered this sentence before.) On being told, "This diagram is not quite right.", GPT-4 produced the following diagram, which I believe is correct except for an extra "s" on "uncle" (you can see this in your browser at https://is.gd/Axufjx):

  digraph G {
      node [shape=plaintext];

      S [label="S"];
      NP [label="NP"];
      VP [label="VP"];

      S -> NP;
      S -> VP;

      // Noun Phrase breakdown
      DP_my [label="DP"];
          DP_my -> My;

      UnclePossessive [label="NP"];
          UnclePossessive -> Uncles;
          UnclePossessive -> ApostropheS;

          NP -> DP_my;
          NP -> UncleCompound;

      UncleCompound [label="NP"];
          SlothSquadResidue[label = "N'"];
              SlothSquadResidue-> Sloth;
              SlothSquadResidue-> SquadResidue;

              SquadResidue[label = "N'"]
                  SquadResidue-> Squad;
                  SquadResidue-> Residue;

              UncleCompound-> UnclePossessive;
              UncleCompound-> SlothSquadResidue;


      // Verb Phrase breakdown
       ModalWould[label = "Modal"];
       ModalWould-> Would;

       AuxHave[label = "Aux"];
       AuxHave-> Have;

       VCompiled[label = "V"];
       VCompiled-> Compiled;

       ObjectNP[label = "NP"];  // Direct object noun phrase
           ArticleA[label = "D"];
           ArticleA-> A;

           MeadowPhrase[label = "N'"];  // 'Meadow' modified by 'periodic'
               AdjPeriodic[label = "Adj"];
               AdjPeriodic-> Periodic; 

               MeadowWord[label = N];
               MeadowWord -> Meadow; 

               MeadowPhrase -> AdjPeriodic; 
               MeadowPhrase -> MeadowWord; 

           PPofNight[label = PP];  // Prepositional phrase 'of night'
               PrepOf[label=P]
                   PrepOf-> Of; 

                   NightNounPhrase[label=N]
                       NightNounPhrase-> Night; 

                   PPofNight -> PrepOf; 
                   PPofNight -> NightNounPhrase; 

           ObjectNP -> ArticleA; 
           ObjectNP -> MeadowPhrase; 
           ObjectNP -> PPofNight;


       VP -> ModalWould;
       VP -> AuxHave;
       VP -> VCompiled
       VP -> ObjectNP

  }
Claude produced a reasonably serviceable diagram on the first try with that prompt:

  digraph SyntacticTree {
    node [shape=plaintext];

    S -> NP1;
    S -> VP;

    NP1 -> DET1 [label="Det"];
    NP1 -> NP2 [label="NP"];

    NP2 -> NP3 [label="NP"];
    NP2 -> N1 [label="N"];

    NP3 -> N2 [label="N"];
    NP3 -> POS [label="Pos"];

    VP -> AUX1 [label="Aux"];
    VP -> AUX2 [label="Aux"];
    VP -> V [label="V"];
    VP -> NP4 [label="NP"];

    NP4 -> DET2 [label="Det"];
    NP4 -> ADJ [label="Adj"];
    NP4 -> N3 [label="N"];
    NP4 -> PP [label="PP"];

    PP -> P [label="P"];
    PP -> NP5 [label="NP"];

    NP5 -> N4 [label="N"];

    DET1 -> "My";
    N2 -> "uncle";
    POS -> "'s";
    N1 -> "sloth";
    N1 -> "squad";
    N1 -> "residue";
    AUX1 -> "would";
    AUX2 -> "have";
    V -> "compiled";
    DET2 -> "a";
    ADJ -> "periodic";
    N3 -> "meadow";
    P -> "of";
    N4 -> "night";
  }
On being told, I think incorrectly, "This diagram is not quite right.", it produced a worse diagram.

So LLMs didn't perform nearly as well on this task as I thought they would, but they also performed much better than you thought they would.

Re: 100x defect tolerance: How we solved the yield problem

#146
post #75

Earlier quoted context omitted.

"While I continue to believe that many people are going to collectively lose trillions of dollars ultimately pursuing "AI" at this stage" Can you please explain more why you think so ? Thank you.

It's a hype cycle with many of the hypers and deciders having zero idea about what AI actually is and how it works. ChatGPT, while amazing, is at its core a token predictor, it cannot ever get to an AGI level that you'd assume to be competitive to a human, even most animals. And just as every other hype cycle, this one will crash down hard. The crypto crashes were bad enough but at least gamers got some very cheap GP…

While I basically agree with everything you say, I have to add some caveats:

ChatGPT, while being as far from true AGI as the Elisa chatbot written in Lisp, is extraordinarily more useful, and being used for many things that previously required humans to write the bullshit, like lobbying and propaganda.

And Crypto... right now BTC is at an historical highest. It could even go higher. And it will eventually crash again. It's the nature of that beast.

Re: 100x defect tolerance: How we solved the yield problem

#147
post #118

Earlier quoted context omitted.

It's a hype cycle with many of the hypers and deciders having zero idea about what AI actually is and how it works. ChatGPT, while amazing, is at its core a token predictor, it cannot ever get to an AGI level that you'd assume to be competitive to a human, even most animals. And just as every other hype cycle, this one will crash down hard. The crypto crashes were bad enough but at least gamers got some very cheap GP…

Those glorified token predictors are the missing piece in the puzzle of general intelligence. There is a long way to go still in putting all those pieces together, but I don't think any of the steps left are in the same order of "we need a miracle breakthrough". That said, I believe that this is going one of two ways: we use AI to make things materially harder for humans, in a scale from "you don't get this job" to "…

I'm sure there are more missing pieces.

We are more than Broca's areas. Our intelligence is much more than linguistic intelligence.

However, and this is also an important point, we have built language models far more capable than any language model a single human brain can have.

Makes me shudder in awe of what's going to happen when we add the missing pieces.

Re: 100x defect tolerance: How we solved the yield problem

#148

Earlier quoted context omitted.

It's a hype cycle with many of the hypers and deciders having zero idea about what AI actually is and how it works. ChatGPT, while amazing, is at its core a token predictor, it cannot ever get to an AGI level that you'd assume to be competitive to a human, even most animals. And just as every other hype cycle, this one will crash down hard. The crypto crashes were bad enough but at least gamers got some very cheap GP…

Why do you think that an AGI can't be a token predictor?

By analogy with human brains: Because our own brains are far more than the Broca's areas in them.

Evolution selects for efficiency.

If token prediction could work for everything, our brains would also do nothing else but token prediction. Even the brains of fishes and insects would work like that.

The human brain has dedicated clusters of neurons for several different cognitive abilities, including face recognition, line detection, body parts self perception, 3D spatial orientation, and so on.

Re: 100x defect tolerance: How we solved the yield problem

#149
post #2

I think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't nee…

Of course many people are going to collectively lose trillions, AI's a very highly hyped industry with people racing into it without an intellectual edge and any temporary achievement by any one company will be quickly replicated and undercut by another using the same tools. Economic success of the individuals swarming on a new technology is not a guarantee whatsoever, nor is it an indicator of the impact of the tech…

[dead]

Re: 100x defect tolerance: How we solved the yield problem

#150

I have a dumb question. Why isn't silicon sold in cubes instead of cylinders?

Crystalline silicon is produced with the Czochralski process (https://en.wikipedia.org/wiki/Czochralski_method), which produces a round ingot. So you'd have to cut away perfectly fine silicon to make something squarish.
Post reply on HN