Live data from Hacker News

AI 2027

ai-2027.com

511–520 of 641 posts

Re: AI 2027

#511
post #496

Earlier quoted context omitted.

> there are some new capabilities that are big, but they are still fundamentally next-token predictors Anthropic recently released research where they saw how when Claude attempted to compose poetry, it didn't simply predict token by token and "react" to when it thought it might need a rhyme and then looked at its context to think of something appropriate, but actually saw several tokens ahead and adjusted for where…

Isn't this just a form of next token prediction? i.e. you'll keep your options open for a potential rhyme if you select words that have many associated rhyming pairs, and you'll further keep your options open if you focus on broad topics over niche

In the same way that human brains are just predicting the next muscle contraction.

Re: AI 2027

#512
Late 2025, "its PhD-level knowledge of every field". I just don't think you're going to get there. There is still a fundamental limitation that you can only be as good as the sources you train on. "PhD-level" is not included in this dataset: in other words, you don't become PhD-level by reading stuff.

Maybe in a few fields, maybe a masters level. But unless we come up with some way to have LLMs actually do original research, peer-review itself, and defend a thesis, it's not going to get to PhD-level.

Re: AI 2027

#513

Earlier quoted context omitted.

You said it right, science fiction. Honestly is exactly the tenor I would expect from the AI hype: this text is completely bereft of any rigour while being dressed up in scientific language. There's no evidence, nothing to support their conclusions, no explanation based on data or facts or supporting evidence. It's purely vibes based. Their promise is unironically "the CEOs of AI companies say AGI is 3 years away" !…

Did you see the supplemental material that explains how they arrived at their timelines/capabilities forecasts? https://ai-2027.com/research

It's not at all clear that performance rises with compute in a linear way, which is what they seem to be predicting. GPT-4.5 isn't really that much smarter than 2023's GPT-4, nor is it at all smarter than DeepSeek.

There might be (strongly) diminishing returns past a certain point.

Most of the growth in AI capabilities has to do with improving the interface and giving them more flexibility. For e.g., uploading PDFs. Further: OpenAI's "deep research" which can browse the web for an hour and summarize publicly-available papers and studies for you. If you ask questions about those studies, though, it's hardly smarter than GPT-4. And it makes a lot of mistakes. It's like a goofy but earnest and hard-working intern.

Re: AI 2027

#514

I think we've actually had capable AIs for long enough now to see that this kind of exponential advance to AGI in 2 years is extremely unlikely. The AI we have today isn't radically different from the AI we had in 2023. They are much better at the thing they are good at, and there are some new capabilities that are big, but they are still fundamentally next-token predictors. They still fail at larger scope longer ter…

> They are much better at the thing they are good at, and there are some new capabilities that are big, but they are still fundamentally next-token predictors. I don't really get this. Are you saying autoregressive LLMs won't qualify as AGI, by definition? What about diffusion models, like Mercury? Does it really matter how inference is done if the result is the same?

> Are you saying autoregressive LLMs won't qualify as AGI, by definition?

No, I am speculating that they will not reach capabilities that qualify them as AGI.

Re: AI 2027

#515

Earlier quoted context omitted.

METR [0] explicitly measures the progress on long term tasks; it's as steep a sigmoid as the other progress at the moment with no inflection yet. As others have pointed out in other threads RLHF has progressed beyond next-token prediction and modern models are modeling concepts [1]. [0] https://metr.org/blog/2025-03-19-measuring-ai-ability-to-com... [1] https://www.anthropic.com/news/tracing-thoughts-language-mod...

At the risk of coming off like a dolt and being super incorrect: I don't put much stock into these metrics when it comes to predicting AGI. Even if the trend of "length of task an AI can reliably do doubles every 7 months" continues, as they say that means we're years away from AI that can complete tasks that take humans weeks or months. I'm skeptical that the doubling trend will continue into that timescale, I think…

There is definitely something qualitatively different about weeks/months long tasks.

It reminds me of the difference between a fresh college graduate and an engineer with 10 years of experience. There are many really smart and talented college graduates.

But, while I am struggling to articulate exactly why, I know that when I was a fresh graduate, despite my talent and ambition, I would have failed miserably at delivering some of the projects that I now routinely deliver over time periods of ~1.5 years.

I think LLM's are really good at emulating the types of things I might say are the types of things that would make someone successful at this if I were to write it down in a couple paragraphs, or an article, or maybe even a book.

But... knowing those things as written by others just would not quite cut it. Learning at those time scales is just very different than what we're good at training LLM's to do.

A college graduate is in many ways infinitely more capable than a LLM. Yet there are a great many tasks that you just can't give an intern if you want them to be successful.

There are at least half a dozen different 1000-page manuals that one must reference to do a bare bones approach at my job. And there are dozens of different constituents, and many thousands of design parameters I must adhere to. Fundamentally, all of these things often are in conflict and it is my job to sort out the conflicts and come up with the best compromise. It's... really hard to do. Knowing what to bend so that other requirements may be kept rock solid, who to negotiate with for different compromises needed, which fights to fight, and what a "good" design looks like between alternatives that all seem to mostly meet the requirements. Its a very complicated chess game where it's hopelessly impossible to brute force but you must see the patterns along the way that will point you like sign posts into a good position in the end game.

The way we currently train LLM's will not get us there.

Until an LLM can take things in it's context window, assess them for importance, dismiss what doesn't work or turns out to be wrong, completely dismiss everything it knows when the right new paradigm comes up, and then permanently alter its decision making by incorporating all of that information in an intelligent way, it just won't be a replacment for a human being.

Re: AI 2027

#516

Earlier quoted context omitted.

Isn't this just a form of next token prediction? i.e. you'll keep your options open for a potential rhyme if you select words that have many associated rhyming pairs, and you'll further keep your options open if you focus on broad topics over niche

In the same way that human brains are just predicting the next muscle contraction.

Except that's not how it works...

Re: AI 2027

#517

Late 2025, "its PhD-level knowledge of every field". I just don't think you're going to get there. There is still a fundamental limitation that you can only be as good as the sources you train on. "PhD-level" is not included in this dataset: in other words, you don't become PhD-level by reading stuff. Maybe in a few fields, maybe a masters level. But unless we come up with some way to have LLMs actually do original r…

> Late 2025, "its PhD-level knowledge of every field". I just don't think you're going to get there.

You think too much of PhDs. They are different. Some of them are just repackaging of existing knowledge. Some are just copy-paste like famous Putin's. Not sure he even rad, to be honest.

Re: AI 2027

#518

I think we've actually had capable AIs for long enough now to see that this kind of exponential advance to AGI in 2 years is extremely unlikely. The AI we have today isn't radically different from the AI we had in 2023. They are much better at the thing they are good at, and there are some new capabilities that are big, but they are still fundamentally next-token predictors. They still fail at larger scope longer ter…

Isn't the brain kind of just a predictor as well, just a more complicated one? Instead of predicting and emitting tokens, we're predicting future outcomes and emitting muscle movements. Which is obviously different in a sense but I don't think you can write off the entire paradigm as a dead end just because the medium is different.

Re: AI 2027

#519
post #459

Earlier quoted context omitted.

You're welcome. I know too many upper middle class educated people that don't want to have kids because they believe the earth will cease to be inhabitable in the next 10 years. It's really bizarre to see and they'll almost certainly regret it when they wake up one day alone in a nursing home, look around and realize that the world still exists. And I think the neuroticism around this topic has led young people into…

I think you're discrediting yourself by talking about dark places and opening your parentheses with anti-depressants. Not all brains function like they're supposed to, people getting help they need shouldn't be stigmatized. You also make no argument about your take on things being the right one, you just oppose their worldview to yours and call theirs wrong like you know it is rather than just you thinking yours is r…

Not sure if you're up on the literature but the chemical imbalance theory of depression has been disproven (or at least no evidence for it).

No one is stigmatizing anything. Just that if you consume doom porn it's likely to affect your attitudes towards life. I think it's a lot healthier to believe you can change your circumstances than to believe you are doomed because you believe you have the wrong brain

https://www.nature.com/articles/s41380-022-01661-0

https://www.quantamagazine.org/the-cause-of-depression-is-pr...

https://www.ucl.ac.uk/news/2022/jul/analysis-depression-prob...

Re: AI 2027

#520
post #496

Earlier quoted context omitted.

> there are some new capabilities that are big, but they are still fundamentally next-token predictors Anthropic recently released research where they saw how when Claude attempted to compose poetry, it didn't simply predict token by token and "react" to when it thought it might need a rhyme and then looked at its context to think of something appropriate, but actually saw several tokens ahead and adjusted for where…

Isn't this just a form of next token prediction? i.e. you'll keep your options open for a potential rhyme if you select words that have many associated rhyming pairs, and you'll further keep your options open if you focus on broad topics over niche

Assuming the task remains just generating tokens, what sort of reasoning or planning would say is the threshold, before it's no longer "just a form of next token prediction?"
Post reply on HN