Live data from Hacker News

The sigmoids won't save you

astralcodexten.com

291–297 of 297 posts

Re: The sigmoids won't save you

#291

Earlier quoted context omitted.

LLM's generate their output words sequentially based on probability (from learned stats). Human's don't operate the same way, the thought happens and then the words are generated to reasonably describe that thought.

What I'm saying is that this is incorrect. An "idea" exists within a model before it generates tokens. This property does not distinguish humans from LLMs. Additionally "from learned stats" doesn't disambiguate between a wider variety of things. I'm not aware of any other way to acquire knowledge from measurements. I'd bet that humans do this differently, based on the fact the humans can get further with less trainin…

> What I'm saying is that this is incorrect. An "idea" exists within a model before it generates tokens.

If that were the case, then the systems would generate words based on the fully resolved idea, but that is not how the LLM systems currently work (per vendors descriptions).

They choose words sequentially and both the specifics of the input as well as the chosen output words significantly impacts not just the rest of the output but the very correctness of the output.

> but not so differently that 'learning stats' would be an inaccurate description.

Agreed, humans are generalizing using some mechanism that can be modeled with math.

But the execution of our reasoning and thought processes is not obviously similar to LLM's next word generation based on probabilities.

Re: The sigmoids won't save you

#292

If you want a model, here's one: LLMs have never demonstrated the ability to go obviously beyond interpolating their training data. It takes an army of paid data producers solving homework problems to give ChatGPT the ability to do your homework. All vibecoded apps that turned out to be successful could put on a geological soil chart with other apps, probably on GitHub somewhere, on the corners. The prediction? They…

See the recent breakthrough "Erdos Problem 1196" which experts couldn't solve for 60 years until ChatGPT Pro did. ChatGPT's key idea was the use of the "Von Mangolt function" which it showed could finally settle the problem. Terry Tao has condensed the AI's proof to around a page. The problem was well-known to experts (in the field of number theory) but it was a Large Language Model that ultimately solved it - which…

[deleted]

Re: The sigmoids won't save you

#293

Earlier quoted context omitted.

What I'm saying is that this is incorrect. An "idea" exists within a model before it generates tokens. This property does not distinguish humans from LLMs. Additionally "from learned stats" doesn't disambiguate between a wider variety of things. I'm not aware of any other way to acquire knowledge from measurements. I'd bet that humans do this differently, based on the fact the humans can get further with less trainin…

> What I'm saying is that this is incorrect. An "idea" exists within a model before it generates tokens. If that were the case, then the systems would generate words based on the fully resolved idea, but that is not how the LLM systems currently work (per vendors descriptions). They choose words sequentially and both the specifics of the input as well as the chosen output words significantly impacts not just the rest…

>that is not how the LLM systems currently work (per vendors descriptions)

Anthropic says of the their model[0]:

"""Claude sometimes thinks in a conceptual space that is shared between languages, suggesting it has a kind of universal “language of thought.”

{...}

Claude will plan what it will say many words ahead, and write to get to that destination. We show this in the realm of poetry, where it thinks of possible rhyming words in advance and writes the next line to get there. This is powerful evidence that even though models are trained to output one word at a time, they may think on much longer horizons to do so."""

Anthropic also created 'golden gate claude'[1] by identifying the region of its architecture that corresponded to the concept of the golden gate bridge and activating it. What would such a region exist for if claude could only think one token at a time?

>the execution of our reasoning and thought processes is not obviously similar to LLM's

"Not obviously similar" I can agree with. I don't think you've identified a way in which they are obviously different, though.

[0] https://www.anthropic.com/research/tracing-thoughts-language...

[1] https://www.anthropic.com/news/golden-gate-claude

Re: The sigmoids won't save you

#294

Earlier quoted context omitted.

We understand it enough to see the obvious massive deficiencies in LLMs. They can predict likely sentences but not evaluate truth or logic. They can fairly reliably record facts about the world but not construct internal models of the world.

> They can predict likely sentences but not evaluate truth or logic. They do probabilistically. So do humans as a matter of fact. The best of us are better at it than LLMs, but that's not persuasive evidence of anything meaningful really. > They can fairly reliably record facts about the world but not construct internal models of the world. You don't know that, unless your presuppose a very specific definition of wor…

Humans do not reason by guessing the next most likely token/word. They use logic, morality and systems of thought they have constructed and shared to help them reason and don’t in any way predict tokens in a sequence - we use words to represent our thoughts and feelings about the world, not to construct them.

You’re constructing a post-hoc fantasy of human thought based on how LLMs work because you are desperate for some reason to believe that they are thinking like humans, but they are not. The process is very different and the results are also different.

Re: The sigmoids won't save you

#295

Earlier quoted context omitted.

Have you used the models, out of interest? They routinely do things autonomously that are not in the training set that would take me 8h, and I wouldn't say I'm slow. The profile of tasks they can do this way is jagged, and maintaining architectural coherence ("months, not hours") is still beyond them, but they're perfectly capable of writing plans and sticking to them.

Yeah, I use them all the time. I just don't see any good argument that it's anything other than statistical pattern matching plus some sort of logic encoded in language. My overfitted LLM obviously didn't arrive at Harry Potter the same way JK Rowling did, so the amount of time she spent writing it is completely irrelevant to any discussion about whether or not the LLM should be able to reproduce it. discussions of A…

I don't think you've addressed the fact that they can do long tasks that aren't in the training set? (And the fact that they're just statistical models isn't very relevant. So am I!)

Re: The sigmoids won't save you

#296

Earlier quoted context omitted.

I think his agenda here is to point out that your probability distribution for AI outcomes should be broad (what you said), but most importantly: this means you must take seriously the possibility that we are gonna get superintelligence quite soon. Basically a lot of people say "but isn't it also pretty likely that we DON'T get superintelligence?" And, yes, it is. But superintelligence being even a remotely plausible…

You want to go to the store to get ice cream. Ice cream is delicious and the value of eating ice cream is a small positive, let's say x. There's a one in ten million chance you'll get hit by a car on the way and die, and your life is infinitely precious, therefore the expected value of going is x times 1 = x, and the one of not going is 1/10m times negative infinity which is negative infinity. You are a rational pers…

There's an objection here when you get to tiny numbers, but surely you wouldn't get ice cream if there was a 10+% chance to get hit by a car?

I think he's saying that >1% or even >10% chunk of your probability mass should be on superintelligence, otherwise you're implicily >99% confident of stalled progress, which seems overconfident. We're not talking about not some infinitesimal fraction here.

Re: The sigmoids won't save you

#297

FYI: The author has predicted that "AGI" will be here in 1-2 years and has staked his public reputation on it. He is personally invested in trendlines being lindy rather than sigmoid. I don't think you can use lindy on trends as if trends are static objects, but that's another conversation.

He only has 1.5 more months. If he's wrong he needs to own it. Same for Eliezer Yudkowsky. But these people have too much riding on their brands. No one has the courage to fess up to being wrong. Given how many podcasts he and others have been on professing this belief, it will be hard to just pretend otherwise.

Yudkowsky has never predicted that "AGI" will be here in 1-2 years. He has been saying frequently for years that it is easier to predict how the AI juggernaut will turn out (i.e., very badly for us) than to predict when the very bad things will happen.

(I don't know about the other guy mentioned above.)

Post reply on HN