Live data from Hacker News

Recent AI model progress feels mostly like bullshit

lesswrong.com

441–450 of 478 posts

Re: Recent AI model progress feels mostly like bullshit

#441

Earlier quoted context omitted.

Massive search overlap though - and some questions (like the golf ball puzzle) can be cached for a long time.

AFAIK they got 15% of unseen queries everyday, so it might be not very simple to design an effective cache layer on that. Semantic-aware clustering of natural language queries and projecting them into a cache-able low rank dimension is a non-trivial problem. Of course, LLM can effectively solve that, but then what's the point of using cache when you need LLM for clustering queries...

Not a search engineer, but wouldn’t a cache lookup to a previous LLM result be faster than a conventional free text search over the indexed websites? Seems like this could save money whilst delivering better results?

Re: Recent AI model progress feels mostly like bullshit

#442

Earlier quoted context omitted.

Yeah I’m a computational biology researcher. I’m working on a novel machine learning approach to inferring cellular behavior. I’m currently stumped why my algorithm won’t converge. So, I describe the mathematics to ChatGPT-o3-mini-high to try to help reason about what’s going on. It was almost completely useless. Like blog-slop “intro to ML” solutions and ideas. It ignores all the mathematical context, and zeros in o…

I tend to prefer Claude over all things ChatGPT so maybe give the latest model a try -- although in some way I feel like 3.7 is a step down from the prior 3.5 model

What do you find inferior in 3.7 compared to 3.5 btw? I only recently started using Claude so I don't have a point of reference.

Re: Recent AI model progress feels mostly like bullshit

#443

Earlier quoted context omitted.

My point wasn't chess specific or that they couldn't have specific training for it. It was a more general "here is something that LLMs clearly aren't being trained for currently, but would also be solvable through reasoning skills" Much in the same way a human who only just learnt the rules but 0 strategy would very, very rarely lose here These companies are shouting that their products are passing incredibly hard ex…

>If an LLM is trained for chess then its performance would just come from memorization, not any kind of "reasoning". If you think you can play chess at that level over that many games and moves with memorization then i don't know what to tell you except that you're wrong. It's not possible so let's just get that out of the way. >These companies are shouting that their products are passing incredibly hard exams, solvi…

> Humans have weird failure modes that are odds with their 'intelligence'. We just choose to call them funny names and laugh about it sometimes. These Machines have theirs. That's all there is to it.

Yes, that's all there is to it and it's not enough. I ain't paying for another defective organism that makes mistakes in entirely novel ways. At least with humans you know how to guide them back on course.

If that's the peak of "AI" evolution today, I am not impressed.

Re: Recent AI model progress feels mostly like bullshit

#444

Earlier quoted context omitted.

>The whole premise on which the immense valuations of these AI companies is based on is that they are learning general reasoning skills from their training on language. And they do, just not always in the ways we expect. >This whole premise crashes and burns if you need task-specific training, like explicit chess training. Everyone needs task specific training. Any human good at chess or anything enough to make it a…

Chess is a very simple game, and having basic general reasoning skills is more than enough to learn how to play it. It's not some advanced mathematics or complicated human interaction - it's a game with 30 or so fixed rules. And chess manuals have numerous examples of actual chess games, it's not like they are pure text talking about the game. So, the fact that LLMs can't learn this sample game despite probably inclu…

As in: they do not have general reasoning skills.

Re: Recent AI model progress feels mostly like bullshit

#445

There are real and obvious improvements in the past few model updates and I'm not sure what the disconnect there is. Maybe it's that I do have PhD level questions to ask them, and they've gotten much better at it. But I suspect that these anecdotes are driven by something else. Perhaps people found a workable prompt strategy by trial and error on an earlier model and it works less well with later models. Or perhaps t…

It is like how I am not impressed by the models when it comes to progress with chemistry knowledge.

Why? Because I know so little about chemistry myself that I wouldn't even know what to start asking the model as to be impressed by the answer.

For the model to be useful at all, I would have to learn basic chemistry myself.

Many though I suspect are in this same situation with all subjects. They really don't know much of anything and are therefore unimpressed by the models response in the same way I am not impressed with chemistry responses.

Re: Recent AI model progress feels mostly like bullshit

#446
post #301
post #40

This is a bit of a meta-comment, but reading through the responses to a post like this is really interesting because it demonstrates how our collective response to this stuff is (a) wildly divergent and (b) entirely anecdote-driven. I have my own opinions, but I can't really say that they're not also based on anecdotes and personal decision-making heuristics. But some of us are going to end up right and some of us ar…

You want to block subjectivity? Write some formulas. There are three questions to consider: a) Have we, without any reasonable doubt, hit a wall for AI development? Emphasis on "reasonable doubt". There is no reasonable doubt that the Earth is roughly spherical. That level of certainty. b) Depending on your answer for (a), the next question to consider is if we the humans have motivations to continue developing AI. c…

A lot of people judge by the lack of their desired outcome. Calling that fear and denial is disingenuous and unfair.

Re: Recent AI model progress feels mostly like bullshit

#447

Earlier quoted context omitted.

Could you share any successful prompting techniques for grounding 3.7, even just a project-specific example?

I use this: I don't want to drastically change my current code, nor do I like being told to create several new files and numerous functions/classes to solve this problem. I want you to think clearly and be focused on the task and don't get wild! I want the most straightforward approach which is elegant, intuitive, and rock solid.

As a caveat, I told it to make minimal code for one task and it completely skipped a super important aspect of it, justifying it by saying that I said "minimal".

Not cool, Claude 3.7, not cool.

Re: Recent AI model progress feels mostly like bullshit

#448

Earlier quoted context omitted.

What are you, an LLM? Look at the results of the first twenty hits and come back, then tell me that they don't speak to that specific issue.

Widely reported does not imply widely known.

How else does an LLM distinguish what is widely known, given there are no statistics collected on the general populations awareness of any given celebrities vices? Robo-apologetics in full force here.

Re: Recent AI model progress feels mostly like bullshit

#449
post #322

Earlier quoted context omitted.

How do you make homework assignments LLM-proof? There may be a huge business opportunity if that actually works, because LLMs are destroying education at a rapid pace.

By giving pen and paper exams and telling your students that the only viable preparation strategy is doing the hw assignments themselves :)

Making in-person tests the only thing that counts toward your grade seems to be a step in the right direction. If students use AI to do their homework, it will only hurt them in the long run.

Re: Recent AI model progress feels mostly like bullshit

#450

Earlier quoted context omitted.

The average human score on USAMO (let alone IMO) is zero, of course. Source: I won medals at Korean Mathematical Olympiad.

I am hesitant to correct a math Olympian, but don't you mean the median?

Average is fine.
Post reply on HN