Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

71–80 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#71
post #56

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Does the as yet unwritten prequel of Idiocracy tell the tale of when we started asking Ai chat bots for facts and this was the point of no return for humanity?

Can you blame the users for asking it, when everyone is selling that as a key defining feature?

I use it for asking - often very niche - questions on advanced probability and simulation modeling, and it often gets those right - why those and not a simple verifiable fact about one of the most popular actors in history?

I don’t know about Idiocracy, but something that I have read specific warnings about is that people will often blame the user for any of the tool’s misgivings.

Re: Recent AI model progress feels mostly like bullshit

#73

There are real and obvious improvements in the past few model updates and I'm not sure what the disconnect there is. Maybe it's that I do have PhD level questions to ask them, and they've gotten much better at it. But I suspect that these anecdotes are driven by something else. Perhaps people found a workable prompt strategy by trial and error on an earlier model and it works less well with later models. Or perhaps t…

In the last year, things like "you are an expert on..." have gotten much less effective in my private tests, while actually describing the problem precisely has gotten better in terms of producing results.

In other words, all the sort of lazy prompt engineering hacks are becoming less effective. Domain expertise is becoming more effective.

Re: Recent AI model progress feels mostly like bullshit

#74

Earlier quoted context omitted.

LLMs will never be good at specific knowledge unless specifically trained for with narrow "if else" statements. Its good for broad general overview such as most popular categories of books in the world.

Really? Open-AI says PhD intelligence is just around the corner!

If we were to survey 100 PhDs how many would know correctly that Paul Newman had an alcohol problem.

Re: Recent AI model progress feels mostly like bullshit

#75
The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions this, but it's ridiculous that these companies never tell us what (if any) efforts have been made to remove test data (IMO, ICPC, etc) from train data.

Re: Recent AI model progress feels mostly like bullshit

#76
post #45

Will LLMs end up like compilers? Compilers are also fundamentally important to modern industrial civilization - but they're not profit centers, they're mostly free and open-source outside a few niche areas. Knowing how to use a compiler effectively to write secure and performative software is still a valuable skill - and LLMs are a valuable tool that can help with that process, especially if the programmer is on the…

This seems like a probable end state, but we're going to have to stop calling LLMs "artificial intelligence" in order to get there.

Yep. I'm looking forward to LLMs/deepnets being considered a standard GOFAI technique with uses and limitations and not "we asked the God we're building to draw us a picture of a gun and then it did and we got scared"

Re: Recent AI model progress feels mostly like bullshit

#77
post #66

Earlier quoted context omitted.

So, in other words, are you saying that AI model progress is the real deal and is not bullshit? That is, as you point out, "all of the models up to o3-mini-high" give an incorrect answer, while other comments say that OpenAIs later models give correct answers, with web citations. So it would seem to follow that "recent AI model progress" actually made a verifiable improvement in this case.

I am pretty sure that they must have meant "up through", not "up to", as the answer from o3-mini-high is also wrong in a way which seems to fit the same description, no?

I tried with 4o and it gave me what I thought was a correct answer:

> Paul Newman was not publicly known for having major problems with alcohol in the way some other celebrities have been. However, he was open about enjoying drinking, particularly beer. He even co-founded a line of food products (Newman’s Own) where profits go to charity, and he once joked that he consumed a lot of the product himself — including beer when it was briefly offered.

> In his later years, Newman did reflect on how he had changed from being more of a heavy drinker in his youth, particularly during his time in the Navy and early acting career, to moderating his habits. But there’s no strong public record of alcohol abuse or addiction problems that significantly affected his career or personal life.

> So while he liked to drink and sometimes joked about it, Paul Newman isn't generally considered someone who had problems with alcohol in the serious sense.

As other's have noted, LLMs are much more likely to be cautious in providing information that could be construed as libel. While Paul Newman may have been an alcoholic, I couldn't find any articles about it being "public" in the same way as others, e.g. with admitted rehab stays.

Re: Recent AI model progress feels mostly like bullshit

#78

LeCun criticized LLM technology recently in a presentation: https://www.youtube.com/watch?v=ETZfkkv6V7Y The accuracy problem won't just go away. Increasing accuracy is only getting more expensive. This sets the limits for useful applications. And casual users might not even care and use LLMs anyway, without reasonable result verification. I fear a future where overall quality is reduced. Not sure how many people / co…

I don't know why anyone is surprised that a statistical model isn't getting 100% accuracy. The fact that statistical models of text are good enough to do anything should be shocking.

Re: Recent AI model progress feels mostly like bullshit

#79
post #42

Earlier quoted context omitted.

Thats not really 'simple' for an LLM. This is a niche information about a specifc person, LLM's train on massive amount of data, the more a topic is being present in the data, the better will the answers be. Also, you can/should use the "research" mode for questions like this.

The question is simple and verifiable - it is impressive to me that it’s not contained in the LLM’s body of knowledge - or rather that it can’t reach the answer. This is niche in the grand scheme of knowledge but Paul Newman is easily one of the biggest actors in history, and the LLM has been trained on a massive corpus that includes references to this. Where is the threshold for topics with enough presence in the da…

[flagged]

Re: Recent AI model progress feels mostly like bullshit

#80
I feel we are already in the era of diminishing returns on LLM improvements. Newer models seem to be more sophisticated implementations of LLM technology + throwing more resources at it, but to me they do not seem fundamentally more intelligent.

I don't think this is a problem though. I think there's a lot of low-hanging fruit when you create sophisticated implementations of relatively dumb LLM models. But that sentiment doesn't generate a lot of clicks.

Post reply on HN