My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…
Recent AI model progress feels mostly like bullshit
171–180 of 478 posts
Re: Recent AI model progress feels mostly like bullshit
#172Earlier quoted context omitted.
Yeah I’m a computational biology researcher. I’m working on a novel machine learning approach to inferring cellular behavior. I’m currently stumped why my algorithm won’t converge. So, I describe the mathematics to ChatGPT-o3-mini-high to try to help reason about what’s going on. It was almost completely useless. Like blog-slop “intro to ML” solutions and ideas. It ignores all the mathematical context, and zeros in o…
It's funny, I have the same problem all the time with typical day to day programming roadblocks that these models are supposed to excel at. I'm talking about any type of bug or unexpected behavior that requires even 5 minutes of deeper analysis. Sometimes when I'm anxious just to get on with my original task, I'll paste the code and output/errors into the LLM and iterate over its solutions, but the experience is like…
Re: Recent AI model progress feels mostly like bullshit
#173Seems like they're looking at how they fail and not considering how they're improving in how they succeed.
The efficiency in DeepSeek's Multi-Head Latent Attention[0] is pure advancement.
Re: Recent AI model progress feels mostly like bullshit
#174The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…
This seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/
Re: Recent AI model progress feels mostly like bullshit
#175This is a bit of a meta-comment, but reading through the responses to a post like this is really interesting because it demonstrates how our collective response to this stuff is (a) wildly divergent and (b) entirely anecdote-driven. I have my own opinions, but I can't really say that they're not also based on anecdotes and personal decision-making heuristics. But some of us are going to end up right and some of us ar…
There is nothing wrong with sharing anecdotal experiences. Reading through anecdotal experiences here can help understand how one's own experience are relatable or not. Moreover, if I have X experience it could help to know if it is because of me doing sth wrong that others have figured out. Furthermore, as we are talking about actual impact of LLMs, as is the point of the article, a bunch of anecdotal experiences ma…
Re: Recent AI model progress feels mostly like bullshit
#176Earlier quoted context omitted.
You missed the end of the supply chain. Paying users. Who magically disappear below market sustaining levels of sales when asked to pay.
> Going from $1M ARR to $100M ARR in 12 months, Cursor is the fastest growing SaaS company of all time Just because it's not reaching the insane hype being pushed doesn't mean it's totally useless
Re: Recent AI model progress feels mostly like bullshit
#177The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…
This seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/
Re: Recent AI model progress feels mostly like bullshit
#178Earlier quoted context omitted.
All of this can be true, and has nothing to do with them having a business model. > NVDA will crash when the AI bubble implodes, > making money, nor will they > They have already hit limiting returns in LLM improvements after staggering investments > and it is clear are nowhere near general intelligence. These are all assumptions and opinions, and have nothing to do with whether or not they have a business model. You…
I consider it a business model if they have plans to make money at some point (no sign of that at openai which are not based on hopium) and are not engaged in fraud like bundling and selling to their own subsidiaries (nvda). These are of course just opinions, I’m not sure we can know facts about such companies except in retrospect.
Re: Recent AI model progress feels mostly like bullshit
#179Earlier quoted context omitted.
> where’s the business model? For who? Nvidia sell GPUs, OpenAI and co sell proprietary models and API access, and the startups resell GPT and Claude with custom prompts. Each one is hoping that the layer above has a breakthrough that makes their current spend viable. If they do, then you don’t want to be left behind, because _everything_ changes. It probably won’t, but it might. That’s the business model
That’s not a business model, it’s a pipe dream. This bubble will be burst by the Trump tariffs and the end of the zirp era. When inflation and a recession hit together hope and dream business models and valuations no longer work.
Re: Recent AI model progress feels mostly like bullshit
#180Earlier quoted context omitted.
I like that it's unmonetized, of course, but that's not why I use AI. I use AI because it's better at search. When I can't remember the right keywords to find something, or when the keywords aren't unique, I frequently find that web search doesn't return what I need and AI does. It's impressive how often AI returns the right answer to vague questions. (not always though)
Google used to return the right answer to vague questions until it decided to return the most lucrative answer to vague questions instead.
Edit: and, more importantly, plenty of people willing to pay a subscription for good quality.