Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

571–580 of 646 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#571

Earlier quoted context omitted.

> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.

> The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. Same story every 4 months and yet still no breakout, winning products. I've been hearing "the AI is good now" and "it 10x's my productivity" for a over a year now. If it were true, why aren't the all-in-AI using companies 10-15 years ahead of their competition yet? Why is it still all buggy, poorly designed…

Right now AI hasn't even managed to replace all the human workers taking orders at the fast food drive thru. That's a job often performed by literal children and companies are still waiting for AI to get good enough for even that. Maybe one day it will be good enough, maybe one day it will outperform humans at such a basic task, but that day is not today. If the hype were anything close to reality, we'd see it everywhere in our lives.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#572

This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…

> 2. those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc. The problem with this angle is that it is still absolutely terrible at doing call center/customer service work, and the profitability story is that the price is going to go up rather than go down. For repetitive physical…

All I ask is 12 monkeys, a library of Shakespeare, and an AI lab full of typewriters. We may not produce a mediocre play, but will IPO as a social media marketing agency.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#573

This is the most grounded and coherent take I’ve seen on the actual realizable value of LLMs.. pretty much since they came out. > the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three: 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc. 2. those who need done a small set of narrowly defined tas…

I think using AI for customer service is really, really ineffective, and it's plain to anybody that has interacted with it. It's a shame that we've had decades of shitty customer service from companies that have the most responsibility and resources to do it, and people in tech have shrugged their shoulders saying "it's unreasonable to expect people to scale up their customer service! They gotta make money!" Now AI c…

I feel that attempts to replace customer service with AI have and will continue to fail. It's not able to do that for both technical reasons (the AI is still to dumb) and business reasons (the AI is usually deliberately scoped to be impotent at helping customers with any important issue).

But, when properly used AI can reduce the customer service load. It can handle a lot of simple questions and status updates. There just needs to be a human available when there's any request that escalates past that simple case.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#574

Earlier quoted context omitted.

> and then everyone will say yeah but playing chess doesn’t mean you’re AGI, it can’t even One, you're not addressing what I wrote above and two, yes, that's absolutely correct. Doing X doesn't qualify something as AGI. If you can't X you can't be AGI. The inverse doesn't hold though. Notably, if you have to retrain the model in order to X then it can't possibly be AGI since if it were _general_ it would be capable o…

Completely arbitrary definition that nobody will agree on, stated as if it’s some self-evident ground truth.

Yes, it is indeed self evident. If it can't figure things out then its intelligence isn't general in which case it can't be AGI by definition.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#575
post #73

The premise in the very first point seems off: > the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers... Even assuming this is how the AI companies are being valued (they're not), the numbers are off. The "value" of most knowledge workers -- based on what enterprises currently pay for th…

There is no chance you can replace all workers without massive second order effects that destroy tons of value.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#576
post #64

Earlier quoted context omitted.

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

I'm just dropping this all over this thread but you're unfortunately mistaken https://dynomight.net/more-chess/

Summarizing here for my dear friends, the guy on the other end managed to fine tune a model gpt-3.5-fine-tune against stockfish vs stockfish games to perform at 1200 elo level against stockfish.

I have been proved wrong I should have quit while I was ahead. /s

Re: Why I'm still bearish on LLMs after Navier-Stokes

#577

Earlier quoted context omitted.

Why does a 6 year old not need any of these guardrails? Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking / influencing about. Maybe not all AGIs have a path to digital singularity. Maybe ou…

Ask a 6 year old to draw a chess board from scratch every turn and they too will make mistakes.

Hm? Drawing the next move's game state based on the current one is a very simple task. Chess academies are probably full of six-year-olds who could do this.

For adults getting into playing tournament chess, losing to some kid barely tall enough to reach the board is a rite of passage.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#578

Earlier quoted context omitted.

When thinking about these valuations, shouldn’t we try to quantify how much knowledge work becomes obsolete if other knowledge workers are automated? I.e. there are a huge amount of knowledge workers employed in businesses that create tools for other knowledge workers. AI won’t automate their work, those businesses will just cease to exist. And then there’s the second order effect: if all the knowledge workers get au…

I think you are committing the lump of labor fallacy [1]. Lots of jobs will disappear, but others will appear. Lots of things (both intellectual and material) that are produced nowadays by humans will be produced in the near future by AI. But humans will be needed to do new things. Take the Hugging Face incident. Why did it happen? Because the people whose task was to set up a testing framework took shortcuts. Why di…

There is no guarantee it will be humans doing the labor.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#579

Earlier quoted context omitted.

huh i didn't even notice, that's how i write all my blog posts too. it looks nicer to me and i don't have to bother checking for "proper" capitalization if everything's just lowercase anyway. didn't realize people struggled to read text that way though, maybe i should change my writing style if this is a common pain point

To me, it reads like a lazy fifth grader wrote it that just couldn't be bothered to press the shift key.

But this is not laziness, it's deliberately done signal being cooler than the rest.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#580
post #479

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

Hi! Thanks for the feedback. I've added an orthography toggle for those who prefer a more conventional look. I will add that I'm not very happy with the readability of my site overall at the moment; if anyone has font or other recommendations for style tweaks to make I'd love to hear them!

OK, I'll be that guy:

Since you are willing to consult LLMs for reviewing your post, why not ask them for feedback on improving the CSS, etc?

Post reply on HN