Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

51–60 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#51
post #6

> So maybe there's no mystery: The AI lab companies are lying, and when they improve benchmark results it's because they have seen the answers before and are writing them down. [...then says maybe not...] Well.. they've been caught again and again red handed doing exactly this. Fool me once shame on you, fool me 100 times shame on me.

Hate to say this but the incentive is growth, not progress. Progress is what enabled the growth, but is also extremely hard to plan and deliver. On the other hand, hype is probably somewhat easier and well-tested approach so no surprise lot of the effort goes into marketing. Markets had repeatedly confirmed that there aren't any significant immediate repercussions for cranking up BS levels in marketing materials, while there are some rewards when it works.

Re: Recent AI model progress feels mostly like bullshit

#52

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

LLMs will never be good at specific knowledge unless specifically trained for with narrow "if else" statements.

Its good for broad general overview such as most popular categories of books in the world.

Re: Recent AI model progress feels mostly like bullshit

#53
post #40

This is a bit of a meta-comment, but reading through the responses to a post like this is really interesting because it demonstrates how our collective response to this stuff is (a) wildly divergent and (b) entirely anecdote-driven. I have my own opinions, but I can't really say that they're not also based on anecdotes and personal decision-making heuristics. But some of us are going to end up right and some of us ar…

Agreed! And with all the gaming of the evals going on, I think we're going to be stuck with anecdotal for some time to come.

I do feel (anecdotally) that models are getting better on every major release, but the gains certainly don't seem evenly distributed.

I am hopeful the coming waves of vertical integration/guardrails/grounding applications will move us away from having to hop between models every few weeks.

Re: Recent AI model progress feels mostly like bullshit

#54
LeCun criticized LLM technology recently in a presentation: https://www.youtube.com/watch?v=ETZfkkv6V7Y

The accuracy problem won't just go away. Increasing accuracy is only getting more expensive. This sets the limits for useful applications. And casual users might not even care and use LLMs anyway, without reasonable result verification. I fear a future where overall quality is reduced. Not sure how many people / companies would accept that. And AI companies are getting too big to fail. Apparently, the US administration does not seem to care when they use LLMs to define tariff policy....

Re: Recent AI model progress feels mostly like bullshit

#55
post #42

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Thats not really 'simple' for an LLM. This is a niche information about a specifc person, LLM's train on massive amount of data, the more a topic is being present in the data, the better will the answers be. Also, you can/should use the "research" mode for questions like this.

The question is simple and verifiable - it is impressive to me that it’s not contained in the LLM’s body of knowledge - or rather that it can’t reach the answer.

This is niche in the grand scheme of knowledge but Paul Newman is easily one of the biggest actors in history, and the LLM has been trained on a massive corpus that includes references to this.

Where is the threshold for topics with enough presence in the data?

Re: Recent AI model progress feels mostly like bullshit

#56

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Does the as yet unwritten prequel of Idiocracy tell the tale of when we started asking Ai chat bots for facts and this was the point of no return for humanity?

Re: Recent AI model progress feels mostly like bullshit

#58
post #42

Earlier quoted context omitted.

Thats not really 'simple' for an LLM. This is a niche information about a specifc person, LLM's train on massive amount of data, the more a topic is being present in the data, the better will the answers be. Also, you can/should use the "research" mode for questions like this.

The question is simple and verifiable - it is impressive to me that it’s not contained in the LLM’s body of knowledge - or rather that it can’t reach the answer. This is niche in the grand scheme of knowledge but Paul Newman is easily one of the biggest actors in history, and the LLM has been trained on a massive corpus that includes references to this. Where is the threshold for topics with enough presence in the da…

The question might be simple and verifiable, but it is not a simple for an LLM to mark a particular question as such. This is the tricky part.

An LLM does not care about your question, it is a bunch of math that will spit out a result based on what you typed in.

Re: Recent AI model progress feels mostly like bullshit

#59

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

So, in other words, are you saying that AI model progress is the real deal and is not bullshit?

That is, as you point out, "all of the models up to o3-mini-high" give an incorrect answer, while other comments say that OpenAIs later models give correct answers, with web citations. So it would seem to follow that "recent AI model progress" actually made a verifiable improvement in this case.

Re: Recent AI model progress feels mostly like bullshit

#60
> Sometimes the founder will apply a cope to the narrative ("We just don't have any PhD level questions to ask")

Please tell me this is not what tech-bros are going around telling each other! Are we implying that the problems in the world, the things that humans collectively work on to maintain the society that took us thousands of years to build up, just aren't hard enough to reach the limits of the AI.

Jesus Christ.

Post reply on HN