Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

131–140 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#131

"are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers, " No, they're really not. They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be. And that they will capture most of that ... which they won't. The Frontier Labs ar…

ai has a >10% chance of causing human extinction, according to anthropic big heads.

if that's true, you are wrong.

if that's false, anthropic is dishonest. why trust a dishonest company to be worth anything?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#132
post #73

The premise in the very first point seems off: > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers... Even assuming this is how the AI companies are being valued (they're not), the numbers are off. The "value" of most knowledge workers -- based on what enterprises currently pay for th…

When thinking about these valuations, shouldn’t we try to quantify how much knowledge work becomes obsolete if other knowledge workers are automated? I.e. there are a huge amount of knowledge workers employed in businesses that create tools for other knowledge workers. AI won’t automate their work, those businesses will just cease to exist. And then there’s the second order effect: if all the knowledge workers get au…

I think you are committing the lump of labor fallacy [1]. Lots of jobs will disappear, but others will appear. Lots of things (both intellectual and material) that are produced nowadays by humans will be produced in the near future by AI. But humans will be needed to do new things.

Take the Hugging Face incident. Why did it happen? Because the people whose task was to set up a testing framework took shortcuts. Why did they? Because there weren't enough people who were assigned to do the job. Why not? Because the job is too new and not enough people are qualified to do it. It's a job that simply did not exist 3 years ago. But 3 years from now, this job might very well employ tens of thousands of high skill knowledge workers.

[1] https://en.wikipedia.org/wiki/Lump_of_labour_fallacy

Re: Why I'm still bearish on LLMs after Navier-Stokes

#133

Not convinced by those points. In particular, I found this very misleading or irrelevant: a typical CPU project anecdotally has about three times as many specification and validation engineers as design engineers and a 5:1 ratio is not unheard of The reason silicon design has such verification to design ratio is because the cost of one bug is many, many orders of magnitude higher than software. Both in dollar cost an…

> The reason ... is because the cost of one bug is many, many orders of magnitude higher than software. Both in dollar cost and in schedule cost (it takes months ... and if you messed up and need to spin a fix, it costs tens of millions of dollars, not counting any design engineering cost). Aren't you just describing waterfall? That's still very prevalent in software engineering, and pretty much any other type of eng…

No. Silicon is on another level. Which is why the EDA verification is an industry on its own.

Sure, there are some software that have similar "can't have bugs" requirements. I imagine the computers on Moon missions also had that kind of high bar. I wouldn't use NASA requirements as a proof for how LLMs should be used.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#134
post #95

> those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc. I have no idea how people can so confidently say that call center work is a “controlled environment” or “repetitive”. It’s almost by definition not repetitive or controlled. Customer support is what I go to when the controll…

Depends on what customer support means.

Typically it means knowledge retrieval from a KB or manipulating a control surface not visible to you.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#135
post #71
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

What levels are they actually at in your experience?

So you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder.

In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#136
the issue most of you seem to not realize is that when you put these models in a loop, you are able to do more and more insane and cool things.

have you guys actually designed, built, and deployed agentic workflows?

it is actually quite hard, requires tons of time spent on evals and testing to ensure accuracy, but when it starts to work it is mind blowing.

there is no going back.

listening to people yap about AI when they have only surface level or one dimensional exposure to LLMs and "AI", but have not actually put innovations to work IN PRACTICE.. is a waste of time

Re: Why I'm still bearish on LLMs after Navier-Stokes

#137

Earlier quoted context omitted.

AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

>LLM can’t beat an avg chess player. Why should that matter?

If something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this.

So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#138
post #81

Earlier quoted context omitted.

So AGI needs to be trained on something to work well on it. Lovely reasoning we have right here. Delusion runs deep in HN circles. I say that as someone heavily invested in AI startups and projects and as someone working in the field. I think most people on HN should touch grass and find real human contact. Lmao Incredible reasoning all around here.

AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count

The fact that LLMs can play chess at any level is a strong indication we are in AGI.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#139
Won’t this change though?

> the present problem of reward hacking can be solved only by rigorous specification by domain experts. the time of domain experts is expensive. rigorous specification is itself a skill, demanding its own expertise outside of a given problem domain. even many skilled software engineers are bad at it. for the vast majority of domains, the intersection of domain experts and specification experts is ludicrously small.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#140
post #65

Earlier quoted context omitted.

It is even worse.. This is a classical reinforcement problem where data generation is easy because the rule set is pre-defined. So you really don't even need any data to start with (but would help).

There are more possible game combinations than atoms in the universe, even those generation of valid game states are as you say pre-defined. that is why models cannot go this route and therefore are poor at chess

Isn’t this exactly how AlphaZero was trained? The rules are known and well defined so the training process can generate games without any outside data.

The only reason LLMs are this bad at chess is because the labs don’t care about chess performance so they’re not going out of their way to train the models for it. The ability they do have is from what chess information happens to be in the training data, plus whatever general reasoning abilities they may be able to apply.

Post reply on HN