Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

151–160 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#151

Earlier quoted context omitted.

US only for now may be the issue? It expands what they had before with AI Overviews, but I’m not sure how new either of those are. It showed up for me organically as an AI Mode tab on a native Google search in Firefox ironically. https://support.google.com/websearch/answer/16011537

Very interesting, thank you!

No worries.

What happens if you go directly to https://google.com/aimode ?

Re: Recent AI model progress feels mostly like bullshit

#152

Earlier quoted context omitted.

Very interesting, thank you!

No worries. What happens if you go directly to https://google.com/aimode ?

It asks me to change some permissions, but that help page says this is only available in the US, so I suppose I'll get blocked right after I change them.

Re: Recent AI model progress feels mostly like bullshit

#153

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

This seems fairly obvious at this point. If they were actually reasoning at all they'd be capable (even if not good) of complex games like chess

Instead they're barely able to eek out wins against a bot that plays completely random moves: https://maxim-saplin.github.io/llm_chess/

Re: Recent AI model progress feels mostly like bullshit

#154

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

This is less an LLM thing than an information retrieval question. If you choose a model and tell it to “Search,” you find citation based analysis that discusses that he indeed had problems with alcohol. I do find it interesting it quibbles whether he was an alcoholic or not - it seems pretty clear from the rest that he was - but regardless. This is indicative of something crucial when placing LLMs into a toolkit. The…

I realise your answer wasn't assertive, but if I heard this from someone actively defending AI it would be a copout. If the selling point is that you can ask these AIs anything then one can't retroactively go "oh but not that" when a particular query doesn't pan out.

Re: Recent AI model progress feels mostly like bullshit

#155
There's some interesting information and analysis to start off this essay, then it ends with:

"These machines will soon become the beating hearts of the society in which we live. The social and political structures they create as they compose and interact with each other will define everything we see around us."

This sounds like an article of faith to me. One could just as easily say they won't become the beating hearts of anything, and instead we'll choose to continue to build a better future for humans, as humans, without relying on an overly-hyped technology rife with error and unethical implications.

Re: Recent AI model progress feels mostly like bullshit

#156

The biggest story in AI was released a few weeks ago but was given little attention: on the recent USAMO, SOTA models scored on average 5% (IIRC, it was some abysmal number). This is despite them supposedly having gotten 50%, 60% etc performance on IMO questions. This massively suggests AI models simply remember the past results, instead of actually solving these questions. I'm incredibly surprised no one mentions th…

Yeah I’m a computational biology researcher. I’m working on a novel machine learning approach to inferring cellular behavior. I’m currently stumped why my algorithm won’t converge.

So, I describe the mathematics to ChatGPT-o3-mini-high to try to help reason about what’s going on. It was almost completely useless. Like blog-slop “intro to ML” solutions and ideas. It ignores all the mathematical context, and zeros in on “doesn’t converge” and suggests that I lower the learning rate. Like, no shit I tried that three weeks ago. No amount of cajoling can get it to meaningfully “reason” about the problem, because it hasn’t seen the problem before. The closest point in latent space is apparently a thousand identical Medium articles about Adam, so I get the statistical average of those.

I can’t stress how frustrating this is, especially with people like Terence Tao saying that these models are like a mediocre grad student. I would really love to have a mediocre (in Terry’s eyes) grad student looking at this, but I can’t seem to elicit that. Instead I get low tier ML blogspam author.

**PS** if anyone read this far (doubtful) and knows about density estimation and wants to help my email is bglazer1@gmail.com

I promise its a fun mathematical puzzle and the biology is pretty wild too

Re: Recent AI model progress feels mostly like bullshit

#157

Earlier quoted context omitted.

NVDA will crash when the AI bubble implodes, and none of those Generative AI companies are actually making money, nor will they. They have already hit limiting returns in LLM improvements after staggering investments and it is clear are nowhere near general intelligence.

All of this can be true, and has nothing to do with them having a business model. > NVDA will crash when the AI bubble implodes, > making money, nor will they > They have already hit limiting returns in LLM improvements after staggering investments > and it is clear are nowhere near general intelligence. These are all assumptions and opinions, and have nothing to do with whether or not they have a business model. You…

I consider it a business model if they have plans to make money at some point (no sign of that at openai which are not based on hopium) and are not engaged in fraud like bundling and selling to their own subsidiaries (nvda).

These are of course just opinions, I’m not sure we can know facts about such companies except in retrospect.

Re: Recent AI model progress feels mostly like bullshit

#158
post #107
post #28

Earlier quoted context omitted.

> where’s the business model? For who? Nvidia sell GPUs, OpenAI and co sell proprietary models and API access, and the startups resell GPT and Claude with custom prompts. Each one is hoping that the layer above has a breakthrough that makes their current spend viable. If they do, then you don’t want to be left behind, because _everything_ changes. It probably won’t, but it might. That’s the business model

You missed the end of the supply chain. Paying users. Who magically disappear below market sustaining levels of sales when asked to pay.

> Going from $1M ARR to $100M ARR in 12 months, Cursor is the fastest growing SaaS company of all time

Just because it's not reaching the insane hype being pushed doesn't mean it's totally useless

Re: Recent AI model progress feels mostly like bullshit

#159
post #116

Earlier quoted context omitted.

There’s a simpler explanation than that’s that the model weights aren’t an information retrieval system and other sequences of tokens are more likely given the totality of training data. This is why for an information retrieval task you use an information retrieval tool similarly to how for driving nails you use a hammer rather than a screw driver. It may very well be you could drive the nail with the screw driver, b…

You think that's a simpler explanation? Ok. I think given the amount of effort that goes into "safety" on these systems that my explanation is vastly more likely than somehow this information got lost in the vector soup despite being attached to his name at the top of every search result[0]. 0 https://www.google.com/search?q=did+paul+newman+have+a+drink...

Except if safety blocked this, it would have also blocked the linked conversation. Alignment definitely distorts behaviors of models, but treating them as information retrieval systems is using a screw driver to drive nails. Your example didn’t refute this.

Re: Recent AI model progress feels mostly like bullshit

#160

My mom told me yesterday that Paul Newman had massive problems with alcohol. I was somewhat skeptical, so this morning I asked ChatGPT a very simple question: "Is Paul Newman known for having had problems with alcohol?" All of the models up to o3-mini-high told me he had no known problems. Here's o3-mini-high's response: "Paul Newman is not widely known for having had problems with alcohol. While he portrayed charact…

Looks like you are using the wrong models

https://g.co/gemini/share/ffa5a7cd6f46

Post reply on HN