Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

211–220 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#211

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Gemini 2.5 Pro

Yes, Paul Newman was known for being a heavy drinker, particularly of beer. 1 He acknowledged his high consumption levels himself. 1. Review: Paul Newman memoir stuns with brutal honesty - AP News

apnews.com

While he maintained an incredibly successful career and public life, accounts and biographies note his significant alcohol intake, often describing it as a functional habit rather than debilitating alcoholism, although the distinction can be debated. He reportedly cut back significantly in his later years.

Re: Recent AI model progress feels mostly like bullshit

#212
post #187

Earlier quoted context omitted.

I just had Cursor Pro + Sonnet 3.7 Max one shot a python script to send this question to every model available through groq. >Found 24 models: llama3-70b-8192, llama-3.2-3b-preview, meta-llama/llama-4-scout-17b-16e-instruct, allam-2-7b, llama-guard-3-8b, qwen-qwq-32b, llama-3.2-1b-preview, playai-tts-arabic, deepseek-r1-distill-llama-70b, llama-3.1-8b-instant, llama3-8b-8192, qwen-2.5-coder-32b, distil-whisper-large-…

I find that everyone who replies with examples like this is an expert using expert skills to get the LLM to perform. Which makes me think why is this a skill that is useful to general public as opposed to another useful skill for technical knowledge workers to add to their tool belt?

I agree. But I will say that at least in my social circles I'm finding that a lot of people outside of tech are using these tools, and almost all of them seem to have a healthy skepticism about the information they get back. The ones that don't will learn one way or the other.

Re: Recent AI model progress feels mostly like bullshit

#213
LLM's are pre-trained to minimize perplexity (PPL), which essentially means that they're trained to model the likelihood distribution of the next words in a sequence.

The amazing thing was that minimizing PPL allowed you to essentially guide the LLM output and if you guided it in the right direction (asked it questions), it would answer them pretty well. Thus, LLMs started to get measured on how well they answered questions.

LLMs aren't trained from the beginning to answer questions or solve problems. They're trained to model word/token sequences.

If you want an LLM that's REALLY good at something specific like solving math problems or finding security bugs, you probably have to fine tune.

Re: Recent AI model progress feels mostly like bullshit

#214

So I guess this was written pre-Gemini 2.5

Meh. I've been using 2.5 with Cline extensively and while it is better it's still an incremental improvement, not something revolutionary. The thing has a 1 million token context window but I can only get a few outputs before I have to tell it AGAIN to stop writing comments.

Are they getting better, definitely. Are we getting close to them performing unsupervised tasks, I don't think so.

Re: Recent AI model progress feels mostly like bullshit

#215

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

oh no. LLMs aren't up on the latest celebrity gossip. whatever shall we do.

Re: Recent AI model progress feels mostly like bullshit

#216

Not sure if its been fixed by now but a few weeks ago I was in the Golden Gate park and wondered if it was bigger than Central park. I asked ChatGPT voice, and although it reported the sizes of the parks correctly (with Golden gate park being the bigger size), it then went and said that Central Park was bigger. I was confused, so Googled and sure enough Golden gate park is bigger. I asked Grok and others as well. I b…

Probably because it has read the facts but has no idea how numbers actually work.

Re: Recent AI model progress feels mostly like bullshit

#217
post #56

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Does the as yet unwritten prequel of Idiocracy tell the tale of when we started asking Ai chat bots for facts and this was the point of no return for humanity?

Some prior works that work as prequels include C.M. Kornbluth's "The Marching Morons" and "The Little Black Bag."

Re: Recent AI model progress feels mostly like bullshit

#218

Earlier quoted context omitted.

> Try to come up with a way to prove humans aren't stochastic parrots Look around you Look at Skyscrapers. Rocket ships. Agriculture. If you want to make a claim that humans are nothing more than stochastic parrots then you need to explain where all of this came from. What were we parroting? Meanwhile all that LLMs do is parrot things that humans created

Skyscrapers: trees, mountains, cliffs, caves in mountainsides, termite mounds, humans knew things could go high, the Colosseum was built two thousand years ago as a huge multi-storey building. Rocket ships: volcanic eruptions show heat and explosive outbursts can fling things high, gunpowder and cannons, bellows showing air moves things. Agriculture: forests, plains, jungle, desert oases, humans knew plants grew from…

You’re likening actual rocketry to LLMs being mildly successful at describing Paul Newman’s alcohol use on average when they already have the entire internet handed to them.

Re: Recent AI model progress feels mostly like bullshit

#219

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Unless you're expecting an LLM to have access to literally all information on earth at all times I find it really hard to care about this particular type of complaint.

My calculator can't conjugate German verbs. That's fine IMO. It's just a tool

Re: Recent AI model progress feels mostly like bullshit

#220

Earlier quoted context omitted.

> Try to come up with a way to prove humans aren't stochastic parrots Look around you Look at Skyscrapers. Rocket ships. Agriculture. If you want to make a claim that humans are nothing more than stochastic parrots then you need to explain where all of this came from. What were we parroting? Meanwhile all that LLMs do is parrot things that humans created

Skyscrapers: trees, mountains, cliffs, caves in mountainsides, termite mounds, humans knew things could go high, the Colosseum was built two thousand years ago as a huge multi-storey building. Rocket ships: volcanic eruptions show heat and explosive outbursts can fling things high, gunpowder and cannons, bellows showing air moves things. Agriculture: forests, plains, jungle, desert oases, humans knew plants grew from…

> when there were over a billion people on Earth in 1800 who could have come up with it

My point is that humans did come up with it. Humans did not parrot it from someone or something else that showed it to us. We didn't "parrot" splitting the atom. We didn't learn how to build skyscrapers from looking at termite hills and we didn't learn to build rockets that can send a person to the moon from seeing a volcano

You are just speaking absolute drivel

Post reply on HN