Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

171–180 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#171

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Perplexity and open-webui+ollama in web search mode answer this question correctly.

Re: Recent AI model progress feels mostly like bullshit

#172

Earlier quoted context omitted.

Yeah I’m a computational biology researcher. I’m working on a novel machine learning approach to inferring cellular behavior. I’m currently stumped why my algorithm won’t converge. So, I describe the mathematics to ChatGPT-o3-mini-high to try to help reason about what’s going on. It was almost completely useless. Like blog-slop “intro to ML” solutions and ideas. It ignores all the mathematical context, and zeros in o…

It's funny, I have the same problem all the time with typical day to day programming roadblocks that these models are supposed to excel at. I'm talking about any type of bug or unexpected behavior that requires even 5 minutes of deeper analysis. Sometimes when I'm anxious just to get on with my original task, I'll paste the code and output/errors into the LLM and iterate over its solutions, but the experience is like…

True. There’s a small bonus that trying to explain the issue to the llm may sometimes be essentially rubber ducking, and that can lead to insights. I feel most of the time the llm can give erroneous output that still might trigger some thinking on a different direction, and sometimes I’m inclined to think it’s helping me more than it actually is.

Re: Recent AI model progress feels mostly like bullshit

#173
> [...] But I would nevertheless like to submit, based off of internal benchmarks, and my own and colleagues' perceptions using these models, that whatever gains these companies are reporting to the public, they are not reflective of economic usefulness or generality. [...]

Seems like they're looking at how they fail and not considering how they're improving in how they succeed.

The efficiency in DeepSeek's Multi-Head Latent Attention[0] is pure advancement.

[0] https://youtu.be/0VLAoVGf_74?si=1YEIHST8yfl2qoGY&t=816

Re: Recent AI model progress feels mostly like bullshit

#174

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

This seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/

Every day I am more convinced that LLM hype is the equivalent of someone seeing a stage magician levitate a table across the stage and assuming this means hovercars must only be a few years away.

Re: Recent AI model progress feels mostly like bullshit

#175
post #40

This is a bit of a meta-comment, but reading through the responses to a post like this is really interesting because it demonstrates how our collective response to this stuff is (a) wildly divergent and (b) entirely anecdote-driven. I have my own opinions, but I can't really say that they're not also based on anecdotes and personal decision-making heuristics. But some of us are going to end up right and some of us ar…

There is nothing wrong with sharing anecdotal experiences. Reading through anecdotal experiences here can help understand how one's own experience are relatable or not. Moreover, if I have X experience it could help to know if it is because of me doing sth wrong that others have figured out. Furthermore, as we are talking about actual impact of LLMs, as is the point of the article, a bunch of anecdotal experiences ma…

Indeed, there’s nothing at all wrong with sharing anecdotes. The problem is when people make broad assumptions and conclusions based solely on personal experience, which unfortunately happens all too often. Doing so is wired into our brains, though, and we have to work very consciously to intercept our survival instincts.

Re: Recent AI model progress feels mostly like bullshit

#176
post #107

Earlier quoted context omitted.

You missed the end of the supply chain. Paying users. Who magically disappear below market sustaining levels of sales when asked to pay.

> Going from $1M ARR to $100M ARR in 12 months, Cursor is the fastest growing SaaS company of all time Just because it's not reaching the insane hype being pushed doesn't mean it's totally useless

I've been here a long time (not this account) and have heard this many times. They all died or became irrelevant.

Re: Recent AI model progress feels mostly like bullshit

#177

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

This seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/

LLMs are capable of playing chess and 3.5 turbo instruct does so quite well (for a human) at 1800 ELO. Does this mean they can truly reason now ?

https://github.com/adamkarvonen/chess_gpt_eval

Re: Recent AI model progress feels mostly like bullshit

#178

Earlier quoted context omitted.

All of this can be true, and has nothing to do with them having a business model. > NVDA will crash when the AI bubble implodes, > making money, nor will they > They have already hit limiting returns in LLM improvements after staggering investments > and it is clear are nowhere near general intelligence. These are all assumptions and opinions, and have nothing to do with whether or not they have a business model. You…

I consider it a business model if they have plans to make money at some point (no sign of that at openai which are not based on hopium) and are not engaged in fraud like bundling and selling to their own subsidiaries (nvda). These are of course just opinions, I’m not sure we can know facts about such companies except in retrospect.

Yep. Facts are usually found out during the SEC investigation but we know that isn't going to happen now...

Re: Recent AI model progress feels mostly like bullshit

#179
post #28

Earlier quoted context omitted.

> where’s the business model? For who? Nvidia sell GPUs, OpenAI and co sell proprietary models and API access, and the startups resell GPT and Claude with custom prompts. Each one is hoping that the layer above has a breakthrough that makes their current spend viable. If they do, then you don’t want to be left behind, because _everything_ changes. It probably won’t, but it might. That’s the business model

That’s not a business model, it’s a pipe dream. This bubble will be burst by the Trump tariffs and the end of the zirp era. When inflation and a recession hit together hope and dream business models and valuations no longer work.

The ZIRP era ended several years ago.

Re: Recent AI model progress feels mostly like bullshit

#180

Earlier quoted context omitted.

I like that it's unmonetized, of course, but that's not why I use AI. I use AI because it's better at search. When I can't remember the right keywords to find something, or when the keywords aren't unique, I frequently find that web search doesn't return what I need and AI does. It's impressive how often AI returns the right answer to vague questions. (not always though)

Google used to return the right answer to vague questions until it decided to return the most lucrative answer to vague questions instead.

Fortunately there is a lot of competition in the LLM space.

Edit: and, more importantly, plenty of people willing to pay a subscription for good quality.

Post reply on HN