Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

141–150 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#141

Earlier quoted context omitted.

Google Labs has AI Mode now, apparently. https://labs.google.com/search/experiment/22

Hm, that's not available to me, what is it? If its an LLM over Google, didn't they release that a few months ago already?

US only for now may be the issue?

It expands what they had before with AI Overviews, but I’m not sure how new either of those are. It showed up for me organically as an AI Mode tab on a native Google search in Firefox ironically.

https://support.google.com/websearch/answer/16011537

Re: Recent AI model progress feels mostly like bullshit

#143

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

And then within a week, Gemini 2.5 was tested and got 25%. Point is AI is getting stronger.

And this only suggested LLMs aren't trained well to write formal math proofs, which is true.

Re: Recent AI model progress feels mostly like bullshit

#144

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

These models are not reliable sources of information. They are either out of date, subject to hallucination, or just plain wrong for a variety of reasons. They are untrustworthy to ask facts like this.

I appreciate your consideration of a subjective question and how you explained it and understand these nuances. But please - do not trust chatgpt etc. I continue to be frustrated at the endless people claiming something is true from chatgpt. I support the conclusions of this author.

Re: Recent AI model progress feels mostly like bullshit

#145
post #112

Earlier quoted context omitted.

It's the first time I've ever used that phrase on HN. Anyway, what phrase do you think works better than 'stochastic parrot' to describe how LLMs function?

Try to come up with a way to prove humans aren't stochastic parrots then maybe people will atart taking you seriously. Just childish reddit angst rn nothing else.

> Try to come up with a way to prove humans aren't stochastic parrots

Look around you

Look at Skyscrapers. Rocket ships. Agriculture.

If you want to make a claim that humans are nothing more than stochastic parrots then you need to explain where all of this came from. What were we parroting?

Meanwhile all that LLMs do is parrot things that humans created

Re: Recent AI model progress feels mostly like bullshit

#146

Earlier quoted context omitted.

Gemini (2.5 Pro): "Yes, Paul Newman was widely known for being a heavy drinker, particularly of beer. He himself acknowledged his significant alcohol consumption." The answer I got ( https://gemini.google.com/share/9e327dc4be03 ) includes references such as https://apnews.com/article/entertainment-reviews-movies-paul... and https://www.psychologytoday.com/us/blog/the-playing-field/20... although they are redacted fro…

Though a local model I'm running (gemma-3-27b-it; https://huggingface.co/lmstudio-community/gemma-3-27b-it-GGU... ) just told me various correct sounding bits about his history with alcohol (correctly citing his alma mater and first wife), but threw in: "Sobriety & AA: Newman got sober in 1964 and remained so for the rest of his life." Which doesn't check out. And it includes plausible but completely hallucinated URL…

Gemma 3 4B (QAT quant): Yes, Paul Newman was indeed known to have struggled with alcohol throughout his life. While he maintained a public image of a charming, clean-cut star, he privately battled alcoholism for many years. He sought treatment in the late 1980s and early 1990s and was reportedly very open about his struggles and the importance of seeking help.

Re: Recent AI model progress feels mostly like bullshit

#147

Earlier quoted context omitted.

Hm, that's not available to me, what is it? If its an LLM over Google, didn't they release that a few months ago already?

US only for now may be the issue? It expands what they had before with AI Overviews, but I’m not sure how new either of those are. It showed up for me organically as an AI Mode tab on a native Google search in Firefox ironically. https://support.google.com/websearch/answer/16011537

Very interesting, thank you!

Re: Recent AI model progress feels mostly like bullshit

#148

The disconnect between improved benchmark results and lack of improvement on real world tasks doesn't have to imply cheating - it's just a reflection of the nature of LLMs, which at the end of the day are just prediction systems - these are language models, not cognitive architectures built for generality. Of course, if you train an LLM heavily on narrow benchmark domains then its prediction performance will improve…

That's fair. But look up the recent experiment on SOTA models on the then just released USAMO 2025 questions. Highest score was 5%, supposedly SOTA last year was IMO silver level. There could be some methodological differences - ie USAMO paper required correct proofs and not just numerical answers. But it really strongly suggests even within limited domains, it's cheating. I'd wager a significant amount that if you t…

> Highest score was 5%, supposedly SOTA last year was IMO silver level.

No LLM last year got silver. Deepmind had a highly specialized AI system earning that

Re: Recent AI model progress feels mostly like bullshit

#149

The core point in this article is that the LLM wants to report _something_, and so it tends to exaggerate. It’s not very good at saying “no” or not as good as a programmer would hope. When you ask it a question, it tends to say yes. So while the LLM arms race is incrementally increasing benchmark scores, those improvements are illusory. The real challenge is that the LLM’s fundamentally want to seem agreeable, and th…

> The real challenge is that the LLM’s fundamentally want to seem agreeable, and that’s not improving

LLMs fundamentally do not want to seem anything

But the companies that are training them and making models available for professional use sure want them to seem agreeable

Re: Recent AI model progress feels mostly like bullshit

#150

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Gemini (2.5 Pro): "Yes, Paul Newman was widely known for being a heavy drinker, particularly of beer. He himself acknowledged his significant alcohol consumption." The answer I got ( https://gemini.google.com/share/9e327dc4be03 ) includes references such as https://apnews.com/article/entertainment-reviews-movies-paul... and https://www.psychologytoday.com/us/blog/the-playing-field/20... although they are redacted fro…

Perplexity:

>Paul Newman is indeed known for having struggled with alcohol during his life. Accounts from various sources, including his own memoir and the documentary ... (https://www.perplexity.ai/search/is-paul-newman-known-for-ha...)

I guess there's something about ChatGPT's set up that makes it different? Maybe they wanted it to avoid libeling people?

Post reply on HN