Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

221–230 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#221
post #219

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Unless you're expecting an LLM to have access to literally all information on earth at all times I find it really hard to care about this particular type of complaint. My calculator can't conjugate German verbs. That's fine IMO. It's just a tool

Yes but a tool for what? When asked a question individuals that don't already have detailed knowledge of a topic are left with no way to tell if the AI generated response is complete bullshit, uselessly superficial, or detailed and on point. The only way to be sure is to then go do the standard search engine grovel looking for authoritative sources.

Re: Recent AI model progress feels mostly like bullshit

#222

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

"Is Paul Newman known for having had problems with alcohol?"

https://chatgpt.com/share/67f332e5-1548-8012-bd76-e18b3f8d52...

Your query indeed answers "...not widely known..."

"Did Paul Newman have problems with alcoholism?"

https://chatgpt.com/share/67f3329a-5118-8012-afd0-97cc4c9b72...

"Yes, Paul Newman was open about having struggled with alcoholism"

What's the issue? Perhaps Paul Newman isn't _famous_ ("known") for struggling with alcoholism. But he did struggle with alcoholism.

Your usage of "known for" isn't incorrect, but it's indeed slightly ambiguous.

Re: Recent AI model progress feels mostly like bullshit

#223

Earlier quoted context omitted.

LLMs are capable of playing chess and 3.5 turbo instruct does so quite well (for a human) at 1800 ELO. Does this mean they can truly reason now ? https://github.com/adamkarvonen/chess_gpt_eval

3.5 turbo instruct is a huge outlier. https://dynomight.substack.com/p/chess Discussion here: https://news.ycombinator.com/item?id=42138289

That might be overstating it, at least if you mean it to be some unreplicable feat. Small models have been trained that play around 1200 to 1300 on the eleuther discord. And there's this grandmaster level transformer - https://arxiv.org/html/2402.04494v1

Open AI, Anthropic and the like simply don't care much about their LLMs playing chess. That or post training is messing things up.

Re: Recent AI model progress feels mostly like bullshit

#224

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

> I use ChatGPT for many tasks every day, but I couldn't fathom that it would get so wrong something so simple.

I think we'll have a term like we have for parents/grandparents that believe everything they see on the internet but specifically for people using LLMs.

Re: Recent AI model progress feels mostly like bullshit

#225

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

"Is Paul Newman known for having had problems with alcohol?" https://chatgpt.com/share/67f332e5-1548-8012-bd76-e18b3f8d52... Your query indeed answers "...not widely known..." "Did Paul Newman have problems with alcoholism?" https://chatgpt.com/share/67f3329a-5118-8012-afd0-97cc4c9b72... "Yes, Paul Newman was open about having struggled with alcoholism" What's the issue? Perhaps Paul Newman isn't _famous_ ("known") f…

Counterpoint: Paul Newman was absolutely a famous drunk, as evidenced by this Wikipedia page.* Any query for "paul newman alcohol" online will return dozens of reputable sources on the topic. Your post is easily interpretable as handwaving apologetics, and it gives big "Its the children who are wrong" energy.

*https://en.wikipedia.org/wiki/Newman_Day

Re: Recent AI model progress feels mostly like bullshit

#226

Earlier quoted context omitted.

This seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/

LLMs are capable of playing chess and 3.5 turbo instruct does so quite well (for a human) at 1800 ELO. Does this mean they can truly reason now ? https://github.com/adamkarvonen/chess_gpt_eval

My point wasn't chess specific or that they couldn't have specific training for it. It was a more general "here is something that LLMs clearly aren't being trained for currently, but would also be solvable through reasoning skills"

Much in the same way a human who only just learnt the rules but 0 strategy would very, very rarely lose here

These companies are shouting that their products are passing incredibly hard exams, solving PHD level questions, and are about to displace humans, and yet they still fail to crush a random-only strategy chess bot? How does this make any sense?

We're on the verge of AGI but there's not even the tiniest spark of general reasoning ability in something they haven't been trained for

"Reasoning" or "Thinking" are marketing terms and nothing more. If an LLM is trained for chess then its performance would just come from memorization, not any kind of "reasoning"

Re: Recent AI model progress feels mostly like bullshit

#227

Earlier quoted context omitted.

3.5 turbo instruct is a huge outlier. https://dynomight.substack.com/p/chess Discussion here: https://news.ycombinator.com/item?id=42138289

That might be overstating it, at least if you mean it to be some unreplicable feat. Small models have been trained that play around 1200 to 1300 on the eleuther discord. And there's this grandmaster level transformer - https://arxiv.org/html/2402.04494v1 Open AI, Anthropic and the like simply don't care much about their LLMs playing chess. That or post training is messing things up.

> That might be overstating it, at least if you mean it to be some unreplicable feat.

I mean, surely there's a reason you decided to mention 3.5 turbo instruct and not.. 3.5 turbo? Or any other model? Even the ones that came after? It's clearly a big outlier, at least when you consider "LLMs" to be a wide selection of recent models.

If you're saying that LLMs/transformer models are capable of being trained to play chess by training on chess data, I agree with you.

I think AstroBen was pointing out that LLMs, despite having the ability to solve some very impressive mathematics and programming tasks, don't seem to generalize their reasoning abilities to a domain like chess. That's surprising, isn't it?

Re: Recent AI model progress feels mostly like bullshit

#228

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

And then within a week, Gemini 2.5 was tested and got 25%. Point is AI is getting stronger. And this only suggested LLMs aren't trained well to write formal math proofs, which is true.

> within a week

How do we know that Gemini 2.5 wasn't specifically trained or fine-tuned with the new questions? I don't buy that a new model could suddenly score 5 times better than the previous state-of-the-art models.

Re: Recent AI model progress feels mostly like bullshit

#229

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

> I use ChatGPT for many tasks every day, but I couldn't fathom that it would get so wrong something so simple. I think we'll have a term like we have for parents/grandparents that believe everything they see on the internet but specifically for people using LLMs.

Look at how many people believe in extremist news outlets!

Re: Recent AI model progress feels mostly like bullshit

#230
post #154

Earlier quoted context omitted.

I realise your answer wasn't assertive, but if I heard this from someone actively defending AI it would be a copout. If the selling point is that you can ask these AIs anything then one can't retroactively go "oh but not that" when a particular query doesn't pan out.

This is a bit of a strawman. There are certainly people who claim that you can ask AIs anything but I don't think the parent commenter ever made that claim. "AI is making incredible progress but still struggles with certain subsets of tasks" is self-consistent position.

It’s not the position of any major AI company, curiously.
Post reply on HN