My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…
Unless you're expecting an LLM to have access to literally all information on earth at all times I find it really hard to care about this particular type of complaint. My calculator can't conjugate German verbs. That's fine IMO. It's just a tool
Recent AI model progress feels mostly like bullshit
221–230 of 478 posts
Re: Recent AI model progress feels mostly like bullshit
#222My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…
https://chatgpt.com/share/67f332e5-1548-8012-bd76-e18b3f8d52...
Your query indeed answers "...not widely known..."
"Did Paul Newman have problems with alcoholism?"
https://chatgpt.com/share/67f3329a-5118-8012-afd0-97cc4c9b72...
"Yes, Paul Newman was open about having struggled with alcoholism"
What's the issue? Perhaps Paul Newman isn't _famous_ ("known") for struggling with alcoholism. But he did struggle with alcoholism.
Your usage of "known for" isn't incorrect, but it's indeed slightly ambiguous.
Re: Recent AI model progress feels mostly like bullshit
#223Earlier quoted context omitted.
LLMs are capable of playing chess and 3.5 turbo instruct does so quite well (for a human) at 1800 ELO. Does this mean they can truly reason now ? https://github.com/adamkarvonen/chess_gpt_eval
3.5 turbo instruct is a huge outlier. https://dynomight.substack.com/p/chess Discussion here: https://news.ycombinator.com/item?id=42138289
Open AI, Anthropic and the like simply don't care much about their LLMs playing chess. That or post training is messing things up.
Re: Recent AI model progress feels mostly like bullshit
#224My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…
I think we'll have a term like we have for parents/grandparents that believe everything they see on the internet but specifically for people using LLMs.
Re: Recent AI model progress feels mostly like bullshit
#225My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…
"Is Paul Newman known for having had problems with alcohol?" https://chatgpt.com/share/67f332e5-1548-8012-bd76-e18b3f8d52... Your query indeed answers "...not widely known..." "Did Paul Newman have problems with alcoholism?" https://chatgpt.com/share/67f3329a-5118-8012-afd0-97cc4c9b72... "Yes, Paul Newman was open about having struggled with alcoholism" What's the issue? Perhaps Paul Newman isn't _famous_ ("known") f…
Re: Recent AI model progress feels mostly like bullshit
#226Earlier quoted context omitted.
This seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/
LLMs are capable of playing chess and 3.5 turbo instruct does so quite well (for a human) at 1800 ELO. Does this mean they can truly reason now ? https://github.com/adamkarvonen/chess_gpt_eval
Much in the same way a human who only just learnt the rules but 0 strategy would very, very rarely lose here
These companies are shouting that their products are passing incredibly hard exams, solving PHD level questions, and are about to displace humans, and yet they still fail to crush a random-only strategy chess bot? How does this make any sense?
We're on the verge of AGI but there's not even the tiniest spark of general reasoning ability in something they haven't been trained for
"Reasoning" or "Thinking" are marketing terms and nothing more. If an LLM is trained for chess then its performance would just come from memorization, not any kind of "reasoning"
Re: Recent AI model progress feels mostly like bullshit
#227Earlier quoted context omitted.
3.5 turbo instruct is a huge outlier. https://dynomight.substack.com/p/chess Discussion here: https://news.ycombinator.com/item?id=42138289
That might be overstating it, at least if you mean it to be some unreplicable feat. Small models have been trained that play around 1200 to 1300 on the eleuther discord. And there's this grandmaster level transformer - https://arxiv.org/html/2402.04494v1 Open AI, Anthropic and the like simply don't care much about their LLMs playing chess. That or post training is messing things up.
I mean, surely there's a reason you decided to mention 3.5 turbo instruct and not.. 3.5 turbo? Or any other model? Even the ones that came after? It's clearly a big outlier, at least when you consider "LLMs" to be a wide selection of recent models.
If you're saying that LLMs/transformer models are capable of being trained to play chess by training on chess data, I agree with you.
I think AstroBen was pointing out that LLMs, despite having the ability to solve some very impressive mathematics and programming tasks, don't seem to generalize their reasoning abilities to a domain like chess. That's surprising, isn't it?
Re: Recent AI model progress feels mostly like bullshit
#228The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…
And then within a week, Gemini 2.5 was tested and got 25%. Point is AI is getting stronger. And this only suggested LLMs aren't trained well to write formal math proofs, which is true.
How do we know that Gemini 2.5 wasn't specifically trained or fine-tuned with the new questions? I don't buy that a new model could suddenly score 5 times better than the previous state-of-the-art models.
Re: Recent AI model progress feels mostly like bullshit
#229My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…
> I use ChatGPT for many tasks every day, but I couldn't fathom that it would get so wrong something so simple. I think we'll have a term like we have for parents/grandparents that believe everything they see on the internet but specifically for people using LLMs.
Re: Recent AI model progress feels mostly like bullshit
#230Earlier quoted context omitted.
I realise your answer wasn't assertive, but if I heard this from someone actively defending AI it would be a copout. If the selling point is that you can ask these AIs anything then one can't retroactively go "oh but not that" when a particular query doesn't pan out.
This is a bit of a strawman. There are certainly people who claim that you can ask AIs anything but I don't think the parent commenter ever made that claim. "AI is making incredible progress but still struggles with certain subsets of tasks" is self-consistent position.