Live data from Hacker News

The changing goalposts of AGI and timelines

mlumiste.com

251–260 of 411 posts

Re: The changing goalposts of AGI and timelines

#251

Earlier quoted context omitted.

A layperson analogy I use is that an LLM is like Dora with a really high IQ - it effectively needs everything reexplained to it, and you can’t give it more than a few seconds of context before it just forgets.

Do you mean Dory, the fish from Finding Nemo?

I have to imagine the poster was referring to Dora the Explorer, a popular and charming cartoon from the start of this century.

Re: The changing goalposts of AGI and timelines

#252

AGI isn't going to happen within the next 30 years so this is moot. The actual researchers have said so many times. It's only the business people and laypeople whooping about AGI always being imminent. You cannot get real, actual AGI (the same ability to perform tasks as a human) without a continuous cycle of learning and deep memory, which LLMs cannot do. The best LLM "memory" is a search engine and document summari…

The post-it note analogy is good, but as a psychiatrist, I'd frame it differently: LLMs are essentially patients with anterograde amnesia.

They can reason brilliantly within a single conversation — just like an amnesic patient can hold an intelligent discussion — but the moment the session ends, everything is gone. No learning happened. No memory formed.

What's worse, even within a session, they degrade. Research shows that effective context utilization drops to This pattern is structurally identical to what I see in clinical practice every day. Anxiety fills working memory with background worry, hallucinations inject noise tokens, depressive rumination creates circular context that blocks updating. In every case, the treatment is the same: clear the context. Medication, sleep, or — for an LLM — a fresh session.

The industry keeps betting on bigger context windows, but that's expanding warehouse floor space while the desk stays the same size. The human brain solved this hundreds of millions of years ago: store everything in long-term memory, recall selectively when needed, consolidate during sleep, and actively forget what's no longer useful.

We can build the smartest single model in the world — the greatest genius humanity has ever seen — but a genius with no memory and no sleep is still just an amnesic savant. The ceiling isn't intelligence. It's architecture.

Re: The changing goalposts of AGI and timelines

#253
post #150

Earlier quoted context omitted.

Turing test is generally misunderstood, much like Schrodinger's cat, it has devolved in to a pop cultural meme. The test is to evaluate if a machine can think . Not if it is intelligent, not if it is human-like. Its dismissed as a useful by most experts in philosophy of mind, AI, language, etc.. Thinking cool and all but not that extraordinary. Even plants does it.

I like the analogy with Schrödinger’s cat. Like Schrödinger’s cat it is actually not a good thought experiment. Both have been debunked. Schrödinger’s cat is applying quantum behavior (of a single interaction) to a macro system (with trillions of interactions). While the Turing test can be explained away with Searle’s Chinese room thought experiment. I would argue that Schrödinger’s cat has done more damage to the ge…

The Turing test and Searle's "rebuttal" are both pretty inconsequential. There's no real definition of "thinking," therefore neither proof/disprove or say much.

Turing's imitation game is about making it difficult for a human to tell whether they are communicating with a computer or not. If a computer can trick the human, then... what? The computer is "thinking" ?

I think most people would say that's an insufficient act to prove thinking. Even though no one has a rigorous definition of thinking either.

All this stuff goes around in circles and like most philosophy makes little progress.

Re: The changing goalposts of AGI and timelines

#254
post #90

Earlier quoted context omitted.

Not in my experience. Quoting my tweet: Gave the same prompt to GPT 5.4 (high) and Opus 4.6 (high). GPT 5.4 implemented the feature, refactored the code (was not asked to), removed comments that were not added in that session, made the code less readable, and introduced a bug. "Undo All". Opus 4.6 correctly recognized that the feature is already implemented in the current code (yeah, lol) and proposed implementing te…

I make ChatGPT and Claude code review each other's outputs. ChatGPT thinks its solutions are better than what Claude produces. What was more surprising to me is that Claude, more often than not, prefers ChatGPT's responses too. I am to sure one can really extrapolate much out of that, but I do find it interesting nonetheless. I think language is also an important factor. I have a hard time deciding which of the two L…

I do the same (I have both review a piece of code), and Codex tend to produce more nitpicky feedback. Opus usually agrees with it on around half the feedback, but says that the other half is too nitpicky to implement. I generally agree with Opus' assessment, and do agree that Codex nitpicks a lot.

I can't even use Codex for planning because it goes down deep design rabbit holes, whereas Opus is great at staying at the proper, high level.

Re: The changing goalposts of AGI and timelines

#255

Anytime I see "Artificial General Intelligence," "AGI," "ASI," etc., I mentally replace it with "something no one has defined meaningfully." Or the long version: "something about which no conclusions can be drawn because the proposed definitions lack sufficient precision and completeness." Or the short versions: "Skippetyboop," "plipnikop," and "zingybang."

They define AGI in their charter > artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work

This definition is not very precise though. For example, I think it can be argued from this definition that we had already reached AGI by the year 2010 (or earlier!). By 2010, computers were integrated into >50% economically valuable work, to the point that humans had mostly forgotten how to do them without computers. Drafting blueprints by hand was already a thing of the past, slide-rules were archaic, paper spreadsheets were long gone. You can debate whether these count as 'highly autonomous', but I don't think it's a clear slam-dunk either way. Not to mention dishwashers, textile weaving machines, CNC machines, assembly lines where >50% is automated, chemical/mineral refining operations, etc.

The definition reminds me of the common quip about robotics, "it's robotics when it doesn't work, once it works it's a machine".

Re: The changing goalposts of AGI and timelines

#256

It's clever and funny, but nobody is legitimately near AGI, and their own AML Corp link proves Altman believes as much: > Achieving AGI, he conceded, will require “a lot of medium-sized breakthroughs. I don’t think we need a big one.” > At the Snowflake Summit in June 2025, Altman predicted that 2026 would mark a breakthrough when AI systems begin generating “novel insights” rather than simply recombining existing in…

We're not near AGI? Personally, I think we've passed it, given that LLMs are now generally more competent than the average person on average.

Re: The changing goalposts of AGI and timelines

#257

Earlier quoted context omitted.

I'm guessing they have a lot of shares in the AI companies they work(ed) for, and they would like to pump their value so they can buy an even nicer carribean island than they can already afford?

Kokotajlo in particular is notable for being the guy who quit OpenAI in 2024 in protest of their policy of requiring researchers to abide by a non-disparagement agreement to retain their equity. In the end OpenAI caved and changed their policy, but if he was lying all along to inflate the value of his shares, it would have been quite a 4d chess move of him to gamble the shares themselves on doing so.

Not quite.

Kokotajlo quit because he didn't think OpenAI would be good stewards of AGI (non-disparagement wasn't in the picture yet). As part of his exit OpenAI asked him to sign a non-disparagement as a condition of keeping his equity. He refused and gave up his equity.

To the best of my knowledge he lost that equity permanently and no longer has any stake in OpenAI (even if this episode later led to an outcry against OpenAI causing them to remove the non-disparagement agreement from future exits).

Re: The changing goalposts of AGI and timelines

#258
post #210

Earlier quoted context omitted.

That’s the problem with the discussions on AI. No one defines the terms they use. If we define AGI as an AI not doing a preset task but can be used for general purpose, then we already have that. If we define it as human level intelligence at _every_ task, then some humans fail to be an AGI. If we define AGI as a magic algorithm that does every task autonomously and successfully then that thing may not exist at all,…

It is not just AGI that is poorly defined. Plain AI is moving goalposts too. When the A* search algorithm was introduced in the late 60s, that was considered AI, when SVM (support vector machines) and KNN (K nearest neighbor) were new, they were AI. And so on. These days it is neural networks and transformer models for language in particular that people mean when they say unqualified AI. It is very hard to have a mea…

I think the Turing test ought to be fine, but we need to be less generous to the AI when executing it. If there exists any human that can consistently tell your AI apart from humans without without insider knowledge, then I don't think you can claim to have AGI. Even if 99.9% of humans can't tell you apart.

So I'm very curious if any AI we have today would pass the Turing test under all circumstances, for example if: the examiner was allowed to continue as long as they wanted (even days/weeks), the examiner could be anybody (not just random selections of humans), observations other than the text itself were fair game (say, typing/response speed, exhaustion, time of day, the examiner themselves taking a break and asking to continue later), both subjects were allowed and expected to search on the internet, etc.

Re: The changing goalposts of AGI and timelines

#259
post #237

Earlier quoted context omitted.

People just overstate their understanding and knowledge, the usual human stuff. The same user has a comment in this thread that contains: 'If you actually know what models are doing under the hood to product output that...' Any one that tells you they know 'what models are dong under the hood' simply has no idea what they're talking about, and it's amazing how common this is.

Fair, I should define what I mean by under the hood. By “under the hood” I mean that models are still just being fed a stream of text (or other tokens in the case of video and audio models), being asked to predict the next token, and then doing that again. There is no technique that anyone has discovered that is different than that, at least not that is in production. If you think there is, and people are just keepin…

If you typed your comment by reading all the others' in the chain, then you responded by typing your response in one go, then you 'just' did next-token prediction based on textual input.

I would still argue that does not prevent you from having intelligence, so that's why this argument is silly.

Re: The changing goalposts of AGI and timelines

#260

Earlier quoted context omitted.

Kokotajlo in particular is notable for being the guy who quit OpenAI in 2024 in protest of their policy of requiring researchers to abide by a non-disparagement agreement to retain their equity. In the end OpenAI caved and changed their policy, but if he was lying all along to inflate the value of his shares, it would have been quite a 4d chess move of him to gamble the shares themselves on doing so.

Isn't it just that he left way before gpt-5, then? At that point a sufficiently naive person could have believed that scaling was going to lead to AGI, but that sort of optimism died after he was already an outsider.

Kokotajlo still believes we get AGI in the next few years. These are his most updated numbers at the moment: https://www.aifuturesmodel.com/
Post reply on HN