Earlier quoted context omitted.
We don’t treat children like they’re stupid, we treat children like they’re children. A stupid adult is treated very differently than any child. Adults are expected to have their world models approximately correct in terms of physical environment so they won’t accidentally kill themselves by falling off a cliff; then there are the social norms which adults are expected to conform to so everyone is kinda predictable t…
You may have been raised properly since you don’t get what I mean. I really envy kids with “Chinese parents” that had them learn math early on and not some bullshit like that if you put your tooth under your pillow, then a tooth fairy will come.
How LLMs work
61–70 of 293 posts
Re: How LLMs work
#62Earlier quoted context omitted.
Yep. It's nearly identical to the neural nets we were using in the 90s. Back then even a supercomputer wasn't big enough or fast enough to do what we do today. I have to wonder though. Is this all a human brain is? A similar thing to an LLM just scaled exponentially larger. I mean a brain is not just neurons with simple connections to each other. The neurons, axons, dendrites, , etc in a brain are all holding and pro…
No, it’s definitely not what a human brain is. That makes very little sense. The ways we interact with language (and thus conceptual memory) is completely and fundamentally different.
If we look beyond written languages which are late inventions of human civilization, oral languages are continuous and build with blocks not words.
Chomskyan school misled the entire field of linguistics for decades by ignoring spoken languages.
Re: How LLMs work
#63Earlier quoted context omitted.
Human brain capabilities are truly amazing, imagine if people didn’t treat their children as if they are stupid and didn’t constantly lie to them, because kids are stupid right, they wouldn’t understand. What heights could be reached.
Because god forbid that childhood, the one time in your life when you don't have any responsibilities, should be fun.
Re: How LLMs work
#64Earlier quoted context omitted.
Could you perhaps cite the core papers for LLMs beyond „Attention is all you need“?
Not a core paper, but I found Formal Algorithms for Transformers [1] (a Google paper from 2022) to have a great pedagogical style. [1] https://arxiv.org/abs/2207.09238
Re: How LLMs work
#65Good article, but when sharing it I will have to preface "yes it's slop, but it's a good explanation".
Absolutely embarrassing that the author didn't catch that these LLM-isms are a (and here I'll use one) bad signal.
In fact, I would go so far as to say that publishing in this style stems from a lack of reading experience and writing experience, which does not bode well for someone pretending to be an expert. I gave this article to someone highly intelligent who doesn't know the first thing about how LLMs work internally, and she immediately called out that it reads like AI text.
Re: How LLMs work
#66Earlier quoted context omitted.
We don’t treat children like they’re stupid, we treat children like they’re children. A stupid adult is treated very differently than any child. Adults are expected to have their world models approximately correct in terms of physical environment so they won’t accidentally kill themselves by falling off a cliff; then there are the social norms which adults are expected to conform to so everyone is kinda predictable t…
You may have been raised properly since you don’t get what I mean. I really envy kids with “Chinese parents” that had them learn math early on and not some bullshit like that if you put your tooth under your pillow, then a tooth fairy will come.
Re: How LLMs work
#67Earlier quoted context omitted.
Those are all just optimizations. We still don’t really know why they work, we just know how to build them.
We do know how they work. They predict the next statistically most likely token. The "bitter lesson" is that fake-it-till-you-make-it is a valid way of doing knowledge work. (Or not make it, then people will just claim you're holding the LLM wrong and it's not the AI's fault.)
Statistically most likely in what context, given which preconditions? Because each prompt sequence is unique so the probability of any token following it is unknown.
Re: How LLMs work
#68Re: How LLMs work
#69Re: How LLMs work
#70Earlier quoted context omitted.
You may have been raised properly since you don’t get what I mean. I really envy kids with “Chinese parents” that had them learn math early on and not some bullshit like that if you put your tooth under your pillow, then a tooth fairy will come.
I think those 2 are orthogonal. Math still works with Santa or the tooth fairy.