Live data from Hacker News

The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

ml-site.cdn-apple.com

271–276 of 276 posts

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#271

Earlier quoted context omitted.

Waymo is a popular argument in self-driving cars, and they do well. However, Waymo is Deep Blue of self-driving cars. Doing very well in a closed space . As a result of this geofencing, they have effectively exhausted their search space, hence they work well as a consequence of lack of surprises. AI works well when search space is limited, but General AI in any category needs to handle a vastly larger search space, a…

This view of Waymo doesn’t account for the fact that self driving is about a lot more than just taking the right roads. It has to deal with other drivers, construction, road closures, pedestrians, bikes, etc.

What I wrote is exactly the opposite. Quoting myself:

> hence they work well as a consequence of lack of surprises. Emphasis mine.

In this context, "lack of surprises" is exactly the rest of the driving besides road choice. In the same space, the behaviors of other actors are also a finite set, or more precisely, can be predicted with much better accuracy.

I drive the same route for ~20 years for commute. The events which surprise me are few and far between, because other people's behavior in that environment is a finite set, and they all behave very predictably, incl. pedestrians, bikes, and other drivers.

Choosing roads are easy, handling surprises hard, but if you saw most potential surprises, then you can drive even without thinking. While I'm not proud of it, my brain took over and drove me home a couple of times on that route when I was too tired to think.

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#272

Earlier quoted context omitted.

Let's outsource it to your mothers house instead.

Let’s do that. Still, this hash table, would be doing reasoning?

The premise of your question is incredibly dumb

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#273
post #252

> Rather than standard benchmarks (e.g., math problems), we adopt controllable puzzle environments that let us vary complexity systematically Very clever, I must say. Kudos to folks who made this particular choice. > we identify three performance regimes: (1) low complexity tasks where standard models surprisingly outperform LRMs, (2) medium-complexity tasks where additional thinking in LRMs demonstrates advantage, a…

Using puzzles is not special or anything, it has been done a million times since before (and including) the LSTM paper (1997) https://www.bioinf.jku.at/publications/older/2604.pdf

The Arc Prize just released a new update and it's all minigame puzzles

https://arcprize.org/

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#275

Earlier quoted context omitted.

FEAR THE ALL-SEEING BASILISK.

Roko's Basilisk has been replaced by Altman's Basilisk. Where once we feared a computer torturing a digital copy of us (Roko's Basilisk), we now fear a computer eliminating all our jobs (Altman's Basilisk). The former has been forgotten, because losing one's job is one step away from losing one's home, which is one of more serious secular deadly sins you can commit in the 21st century. I wait with baited breathe to s…

"bated breath", dammit!

- an old fisherman and aficionado of William Shakespeare.

https://www.vocabulary.com/articles/pardon-the-expression/ba...

FTFA: "Unless you've devoured several cans of sardines in the hopes that your fishy breath will lure a nice big trout out of the river, baited breath is incorrect."*

Re: The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

#276
The worse piece of news is that the performance of thinking and non-thinking models is inverted depending on whether the problem is simple or medium, which implies that you need to know the problem's complexity to choose the best model. Ouch.
Post reply on HN