Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

481–490 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#481

This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…

In my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law (https://en.wikipedia.org/wiki/Parkinson%27s_Law) is still in effect. You could even expand "Work expands so as to fill the time available for its completion" with "time and tokens available".

Re: Why I'm still bearish on LLMs after Navier-Stokes

#482

Earlier quoted context omitted.

Do you want to measure the ability of the box, or measure the ability of the box with one hand tied behind its back? More to my point, I think it's stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs. I'm advising people that they should think about this distinction, themselves, when they have data and want answers.

Neither. As I said I want to measure cognitive abilities. Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.

I agree that it's a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had.

That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way.

Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce.

Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: https://www.youtube.com/watch?v=gDy9AUQJ3Fg

This lamp, without a working power outlet? It really doesn't do anything...

Re: Why I'm still bearish on LLMs after Navier-Stokes

#483
post #3

I really appreciate seeing a tempered take that's not literally denialist about current capabilities.

I don’t know who you’re talking about, even the most bearish people like Gary Marcus and Ed Zitron acknowledge that LLMs are useful in these same cases the OP admits. Gary Marcus is even still a long term AI advocate, he just doesn’t think LLMs are enough and we need more foundational breakthroughs. Zitron says it’s valuable technology but not worth the trillion dollar valuations the frontier labs are claiming. The l…

There are some people who call literally anything crated with the assistance of AI “slop”. Doesn’t matter how or to what extent, it’s all slop from the slop machine to them.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#484

Earlier quoted context omitted.

What happens when you ask those same frontier models to write a chess-playing program? I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those func…

>What happens when you ask those same frontier models to write a chess-playing program? they shit out a carbon copy of https://github.com/official-stockfish/stockfish that they have in their training data. Still doesn't make Fable good at playing chess.

I feel like you're saying something as odd as "Transistors still aren't good at playing chess."

I'm pretty sure Fable could write AlphaZero, which has no lineage in common with stockfish.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#485

Earlier quoted context omitted.

I tested both myself and a weak bot against Astra xhigh, https://lichess.org/study/27lCQqDa . It's still pretty bad at chess, though it takes longer to devolve into illegal moves.

> though it takes longer to devolve into illegal moves Is this because the context is being saturated? How did you set it up? Was the prompt something like "Here's the state of the board, you're white, your move, what do you do?" and then starting fresh each time? Or did it include the whole history of moves and board states and previous thinking tokens and so on? No judgment, just trying to add this data point (than…

You can see the entire conversation for my game at https://chatgpt.com/share/6aaac17b-1384-83e8-98fd-4350a0ef69....

Re: Why I'm still bearish on LLMs after Navier-Stokes

#486

Earlier quoted context omitted.

Neither. As I said I want to measure cognitive abilities. Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.

I agree that it's a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had. That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way. Like, to judge their utility, I think we should be open to letting them…

I completely agree.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#488
post #95

> those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc. I have no idea how people can so confidently say that call center work is a “controlled environment” or “repetitive”. It’s almost by definition not repetitive or controlled. Customer support is what I go to when the controll…

It's repetitive and controlled if you don't care about the outcome, which monopoly companies don't.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#489

This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…

In my own experience as a software developer, I see no justification for the hype. Demand for software is practically infinite, and agents aren't mind readers so they'll always need workers to turn needs into prompts. Despite how much work they get done on their own, Parkinson's Law ( https://en.wikipedia.org/wiki/Parkinson%27s_Law ) is still in effect. You could even expand "Work expands so as to fill the time avail…

Why do you belive that agents wouldn't be able to take over product management, and generate prompts for the "software engineer" agents?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#490

The challenge with estimating abilities, is that we don’t know what the models can achieve if we just burn enough money. The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something. It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem? IMO the very best case scenario / potential for th…

> The navier-stokes shows us what mathematical problem can be solved when $10m worth of compute is thrown at something. > It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem? We need to have robotics automation catchup first. The math and coding problems are problems in written-space only: you can set up feedback loops to test what worked a…

You just invented Skynet.
Post reply on HN