Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

351–360 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#351

I asked Grok to review the comments here and generate a response defending AI: After reviewing the discussion on the Hacker News thread, it’s clear that there are a range of complaints and criticisms about AI, particularly centered around its limitations, overhype, and practical utility. Some users express frustration with AI’s inability to handle complex reasoning, its tendency to produce generic or incorrect output…

[dead]

Re: Recent AI model progress feels mostly like bullshit

#352

Earlier quoted context omitted.

There is nothing wrong with sharing anecdotal experiences. Reading through anecdotal experiences here can help understand how one's own experience are relatable or not. Moreover, if I have X experience it could help to know if it is because of me doing sth wrong that others have figured out. Furthermore, as we are talking about actual impact of LLMs, as is the point of the article, a bunch of anecdotal experiences ma…

Indeed, there’s nothing at all wrong with sharing anecdotes. The problem is when people make broad assumptions and conclusions based solely on personal experience, which unfortunately happens all too often. Doing so is wired into our brains, though, and we have to work very consciously to intercept our survival instincts.

People "make conclusions" because they have to take decisions day to day. We cannot wait for the perfect bulletproof evidence before that. Data is useful to take into account, but if I try to use X llm that has some perfect objective benchmark backing it, while I cannot make it be useful to me while Y llm has better results, it would be stupid not to base my decision on my anecdotal experience. Or vice versa, if I have a great workflow with llms, it may be not make sense to drop it because some others may think that llms don't work.

In the absence of actually good evidence, anecdotal data may be the best we can get now. The point imo is try to understand why some anecdotes are contrasting each other, which, imo, is mostly due to contextual factors that may not be very clear, and to be flexible enough to change priors/conclusions when something changes in the current situation.

Re: Recent AI model progress feels mostly like bullshit

#353
post #345

Earlier quoted context omitted.

Yes, here's the link: https://arxiv.org/abs/2503.21934v1 Anecdotally, I've been playing around with o3-mini on undergraduate math questions: it is much better at "plug-and-chug" proofs than GPT-4, but those problems aren't independently interesting, they are explicitly pedagogical. For anything requiring insight, it's either: 1) A very good answer that reveals the LLM has seen the problem before (e.g. naming the theo…

This is a paper by INSAIT researchers - a very young institute which hired most of its PHD staff only in the last 2 years, basically onboarding anyone who wanted to be part of it. They were waiving their BG-GPT on national TV in the country as a major breakthrough, while it was basically was a Mistral fine-tuned model, that was eventually never released to the public, nor the training set. Not sure whether their (INS…

[dead]

Re: Recent AI model progress feels mostly like bullshit

#354

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

Less than 5%. OpenAI's O1 burned through over $100 in tokens during the test as well!

Re: Recent AI model progress feels mostly like bullshit

#355

Earlier quoted context omitted.

What would the average human score be? I.e. if you randomly sampled N humans to take those tests.

The average human score on USAMO (let alone IMO) is zero, of course. Source: I won medals at Korean Mathematical Olympiad.

I am hesitant to correct a math Olympian, but don't you mean the median?

Re: Recent AI model progress feels mostly like bullshit

#356

Earlier quoted context omitted.

> That might be overstating it, at least if you mean it to be some unreplicable feat. I mean, surely there's a reason you decided to mention 3.5 turbo instruct and not.. 3.5 turbo? Or any other model? Even the ones that came after? It's clearly a big outlier, at least when you consider "LLMs" to be a wide selection of recent models. If you're saying that LLMs/transformer models are capable of being trained to play ch…

I mentioned it because it's the best example. One example is enough to disprove the "not capable of". There are other examples too. >I think AstroBen was pointing out that LLMs, despite having the ability to solve some very impressive mathematics and programming tasks, don't seem to generalize their reasoning abilities to a domain like chess. That's surprising, isn't it? Not really. The LLMs play chess like they have…

The issue isn’t whether they can be trained to play. The issue is whether, after making a careful reading of the rules, they can infer how to play. The latter is something a human child could do, but it is completely beyond an LLM.

Re: Recent AI model progress feels mostly like bullshit

#357

Earlier quoted context omitted.

"Paul Newman alcohol" is just showing you results where those words are all present, it's not really implying how widely known it is.

What are you, an LLM? Look at the results of the first twenty hits and come back, then tell me that they don't speak to that specific issue.

Widely reported does not imply widely known.

Re: Recent AI model progress feels mostly like bullshit

#358

Earlier quoted context omitted.

Skyscrapers: trees, mountains, cliffs, caves in mountainsides, termite mounds, humans knew things could go high, the Colosseum was built two thousand years ago as a huge multi-storey building. Rocket ships: volcanic eruptions show heat and explosive outbursts can fling things high, gunpowder and cannons, bellows showing air moves things. Agriculture: forests, plains, jungle, desert oases, humans knew plants grew from…

> when there were over a billion people on Earth in 1800 who could have come up with it My point is that humans did come up with it. Humans did not parrot it from someone or something else that showed it to us. We didn't "parrot" splitting the atom. We didn't learn how to build skyscrapers from looking at termite hills and we didn't learn to build rockets that can send a person to the moon from seeing a volcano You a…

It's obvious that humans imitate concepts and don't come up with things de-novo from a blank slate of pure intelligence. So your claim hinges on LLMs parrotting the words they are trained on. But they don't do that, their training makes them abstract over concepts and remix them in new ways to output sentences they weren't trained on, e.g.:

Prompt: "Can you give me a URL with some novel components, please?"

DuckDuckGo LLM returns: "Sure! Here’s a fictional URL with some novel components: https://www.example-novels.com/2023/unique-tales/whimsical-j..."

An living parrot echoing "pieces of eight" cannot do this, it cannot say "pieces of " or "pieces of " even if asked to do that. The LLM training has abstracted some concept of what it means for a text pattern to be a URL and what it means for things to be "novel" and what it means to switch out the components of a URL but keep them individually valid. It can also give a reasonable answer asking for a new kind of protocol. So your position hinges on the word "stochastic" which is used as a slur to mean "the LLM isn't innovating like we do it's just a dice roll of remixing parts it was taught". But if you are arguing that makes it a "stochastic parrot" then you need to consider splitting the atom in its wider context...

> "We didn't "parrot" splitting the atom"

That's because we didn't "split the atom" in one blank-slate experiment with no surrounding context. Rutherford and team disintegrated the atom in 1914-1919 ish, they were building on the surrounding scientific work happening at that time: 1869 Johann Hittorf recognising that there was something coming in a straight line from or near the cathode of a Crookes vacuum tube, 1876 Eugen Goldstein proving they were coming from the cathode and naming them cathode rays (see: Cathode Ray Tube computer monitors), and 1897 J.J Thompson proving the rays are much lighter than the lightest known element and naming them Electrons, the first proof of sub-atomic particles existing. He proposed the model of the atom as a 'plum pudding' (concept parroting). Hey guess who JJ Thomspon was an academic advisor of? Ernest Rutherford! 1911 Rutherford discovery of the atomic nucleus. 1909 Rutherford demonstrated sub-atomic scattering and Millikan determined the charge on an electron. Eugen Goldstein also discovered the anode rays travelling the other way in the Crookes tube and that was picked up by Wilhelm Wien and it became Mass Spectrometry for identifying elements. In 1887 Heinrich Hertz was investigating the Photoelectric effect building on the work of Alexandre Becquerel, Johann Elster, Hans Geitel. Dalton's atomic theory of 1803.

Not to mention Rutherford's 1899 studies of radioactivity, following Henri Becquerel's work on Uranium, following Marie Curie's work on Radium and her suggestion of radioactivity being atoms breaking up, and Rutherford's student Frederick Soddy and his work on Radon, and Paul Villard's work on Gamma Ray emissions from Radon.

When Philipp Lenard was studying cathode rays in the 1890s he bought up all the supply of one phosphorescent material which meant Röntgen had to buy a different one to reproduce the results and bought one which responded to X-Rays as well, and that's how he discovered them - not by pure blank-sheet intelligence but by probability and randomness applied to an earlier concept.

That is, nobody taught humans to split the atom and then humans literally parotted the mechanism and did it, but you attempting to present splitting the atom as a thing which appeared out of nowhere and not remixing any existing concepts is, in your terms, absolute drivel. Literally a hundred years and more of scientists and engineers investigating the subatomic world and proposing that atoms could be split, and trying to work out what's in them by small varyations on the ideas and equipment and experiments seen before, you can just find names and names and names on Wikipedia of people working on this stuff and being inspired by others' work and remixing the concepts in it, and we all know the 'science progresses one death at a time' idea that individual people pick up what they learned and stick with it until they die, and new ideas and progress need new people to do variations on the ideas which exist.

No people didn't learn to build rockets from "seeing a volcano" but if you think there was no inspiration from fireworks, cannons, jellyfish squeezing water out to accelerate, no sudies of orbits from moons and planets, no chemistry experiments, no inspiration from thousands of years of flamethrowers: https://en.wikipedia.org/wiki/Flamethrower#History no seeing explosions moving large things, you're living in a dream

Re: Recent AI model progress feels mostly like bullshit

#359

Earlier quoted context omitted.

> The real challenge is that the LLM’s fundamentally want to seem agreeable, and that’s not improving LLMs fundamentally do not want to seem anything But the companies that are training them and making models available for professional use sure want them to seem agreeable

That sound reasonable to me, but the those companies forget that there's different types of agreeable. There's the LLM approach, similar to the coworker who will answer all your questions about .NET but not stop you from coding yourself into a corner, and then there's the "Let's sit down and review what it actually is that you're doing, because you're asking a fairly large number of disjoint questions right now". I'v…

I think it's your responsibility to control the LLM. Sometimes, I worry that I'm beginning to code myself into a corner, and I ask if this is the dumbest idea it's ever heard and it says there might be a better way to do it. Sometimes I'm totally sceptical and ask that question first thing. (Usually it hallucinates when I'm being really obtuse though, and in a bad case that's the first time I notice it.)

Re: Recent AI model progress feels mostly like bullshit

#360

Earlier quoted context omitted.

> The real challenge is that the LLM’s fundamentally want to seem agreeable, and that’s not improving LLMs fundamentally do not want to seem anything But the companies that are training them and making models available for professional use sure want them to seem agreeable

> LLMs fundamentally do not want to seem anything You're right that LLMs don't actually want anything. That said, in reinforcement learning, it's common to describe models as wanting things because they're trained to maximize rewards. It’s just a standard way of talking, not a claim about real agency.

Reinforcement learning, maximise rewards? They work because rabbits like carrots. What does an LLM want? Haven't we already committed the fundamental error when we're saying we're using reinforcement learning and they want rewards?
Post reply on HN