Earlier quoted context omitted.
fair reply to an hour video, Scott is just so good to hear his talk is better than I can explain it... go to 24 minutes and 07 seconds. it's statistically determining what the next word should be based on all the text it's been trained on. It's not intelligence and he shows what probability it puts on each word that it chooses, but also shows a lot of the other words it was thinking of using. In a later part he shows…
Ok, I want to thank you for finally giving us a concrete falsifiable statement that we can check. I pretended Marseille 40 times before asking Luna 5.6, and the answer was Paris. So, even with concrete examples, model haters are still wrong. You also imply the claim that making the distribution of words as the possible next one visible, somehow makes the whole system not intelligent. I would say the exact opposite is…
I don't see the many weighted words as a weakness, I see it opening up what's under the hood of the prediction machine that it is.
LLMs are very cool tech, definitely not a model hater, the use case on when to use it makes a difference, it's not AGI.