Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

631–640 of 646 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#631

Earlier quoted context omitted.

yes, if you use a tool improperly that's on you. i am specifically talking about people who don't use skills (the most basic shit) and then get sloppy code.

But what if I use the skills, and still can't model to follow it? Does the AI company say anywhere that they will refund the tokens if the model does not follow what is written in the Skills file? If there is no such guarantee, why should I spend time writing an elaborate skill file? There is no telling when the model chose to ignore stuff in it. I don't understand how people can work with something like that!

i think you miss the part about the loop. you're still building software for the past/current world.

when you have chains of agents working seamlessly, where agents are managed by themselves (like spawning a new chain to do some scope of work), and it just sort of works on a infinite game loop, your system continually keeps improving as it works. at least, that is the goal with the systems i like to build.

you have to do your due diligence, obviously, as you would with any TOOL you use. in that case it means having specific goals, methodologies, etc. that each model has to follow. (if one is a "hey grok explain this" type of person when it comes to using ai, ngmi)

we have one agentic workflow where there are 4 different models that can be spawned (to handle tasks of various complexity), that work against an API (the source of truth). their goal is to work any time a specific file is uploaded, and handle it. it is a complicated file, with 1000s of line items, with varying amounts of uncertainty involved.

it happens in the real world today, where it costs $xxx,xxx per year to do. because it involves many people and companies... it is done totally manually today..

how it works in order to get to that resolution is different based on the complexity of the task... sometimes there's bad data in the mix because people make mistakes when they create stuff (we're not special). in this case it requires finding out whether it is indeed bad data or actually correct. that requires setting up a scheduled job, and handling it when there is a resolution.

i don't code this part. how this gets accomplished... the agents are able to work autonomously to handle any edge case, in order to complete the goal. the goal in this case is to ultimately record transactions (a goal of a business is to make money, believe it or not).

does this not make sense? jeez hacker news used to be an imaginative place.

the model that we pay $10-20 per 1M tokens will be $1 per 1M token next year.

And the model next year that we can pay $10-20 per 1M token will be even better than Astra and Terra/Sol/Luna. It is an exciting time.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#632

Earlier quoted context omitted.

lol. wish it was. fwiw i was pretty much bearish on AI for years because I did go deep into this shit since i got access to GPT 3. but since astra i have been quite bullish. 98% of people think AI means what gemini tells them when they do a google search. of that 2% who go beyond... maybe 20% of those are using AI to code. most SWE still think "using AI" to code means the copilot pane they open on the side in VSC. th…

> 98% of people think AI means what gemini tells them when they do a google search. > of that 2% who go beyond... maybe 20% of those are using AI to code. > so of that 20%, maybe 5% have... maybe 1-5% of that 5% actually went deeper How can you write any of this when you said you just started liking LLMs with Astra, a model that released two weeks ago. You're whipping out a bunch of made up stats trying to describe l…

because astra was a big jump, and it is a new model. i didn't get these types of results with fable. been building agentic workflows since gpt 3 days, and using AI coding agents since 2023.

been building stuff for years, so i'm not just some w2/1099 type guy who bike sheds for a living lol

this stuff is new... obviously it is a fast moving space.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#633
post #479

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

Hi! Thanks for the feedback. I've added an orthography toggle for those who prefer a more conventional look. I will add that I'm not very happy with the readability of my site overall at the moment; if anyone has font or other recommendations for style tweaks to make I'd love to hear them!

FWIW, I too found it harder to read without capitalization. Came here, read through comments, saw this comment, thought "Oh I didn't see a toggle, did they just add it literally while I was reading?". Went back to the site and realized the toggle was the "orthography mode: based|cringe" bit, which I would have never realized would do that.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#634

Earlier quoted context omitted.

I feel like you're saying something as odd as "Transistors still aren't good at playing chess." I'm pretty sure Fable could write AlphaZero, which has no lineage in common with stockfish.

I'm not the one posting daily about how "AIs are going to destroy the world because of how smart they are", "humans are finished" and "we've reached super duper mega intelligence". Go see Dario and Sam about that. >I'm pretty sure Fable could write AlphaZero If course it does, the paper is open and dozens of open source implementations are in its training data already. It could write AlphaStockfish, or xx_chessmaster…

So if you need a scratchpad to solve a problem, does that mean you cannot solve that problem?

Re: Why I'm still bearish on LLMs after Navier-Stokes

#635

Earlier quoted context omitted.

Throughout this exchange you're repeatedly confusing the negative and the positive. I agree with you that there is no rigorous and universally agreed upon criteria for exactly what would constitute AGI (ie the positive). There are some vague shapes that are widely (but not universally) accepted such as largely (vague boundary) being capable of replacing (vague criteria) humans. However there are plenty of disqualifie…

> However there are plenty of disqualifiers that are more or less universally accepted (ie the negative) Which is exactly the point I’ve made repeatedly, there will always be something that they cannot do, and thus there will never be AGI . There will always be a long tail of capabilities that whatever system is created doesn’t have, and a long line of social media commenters eager to list them. An AI controlled robo…

[dead]

Re: Why I'm still bearish on LLMs after Navier-Stokes

#636
post #344

Earlier quoted context omitted.

It's an interesting puzzle, isn't it. On the one hand, the AIs are no good at playing Chess. However, on the other hand, if you ask an AI to win a game of chess it has all the tools on hand to compete at the same level as Stockfish - it can re-implement an engine and even probably has a GPU on hand to train its own neural nets. So should we say that the AI can play chess well, or that it cannot?

Is it, though? If you design and build a winning F1 race car, did you also win the race? Recent discourse around AI seems to conflate the semantics of winning: 1. you contributed to the win vs 2. you yourself were the winning driver.

But the human would have to get extra hardware to do that. The AI isn't bringing in any resources it doesn't already have access to.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#637
post #555

Earlier quoted context omitted.

IMO the fear is that software engineering becomes a _high_-skill profession. If the models get good enough to solve all of the low-skill problems, then we only need to keep around the people who are highly skilled. That means a fewer number of software engineers, and a difficult path to becoming someone who is highly skilled.

It's already a fairly high skill profession in general. The seemingly-lower-skilled roles tend to be the ones with higher design/creative requirements on the programmers. Agreed about the barrier to entry raising though, we're already seeing that in the glut of CS grads who can't actually get a programming job right now. Personally I suspect that this is a cultural issue more than an economic one, though. Companies n…

Maybe I'm misunderstanding you but I believe I'd disagree with your first point on the perceived lower skill roles with higher design / creative requirements. I've seen a lot of the software industry being software factories rather than something closer to craftsmanship or even proper ABET-style engineering and AI is indeed better than the manually written slop that was encouraged in low-creativity / low-innovation environments such as QA and a lot of infrastructure engineering like defining CI pipelines. It seems reminiscent of the waves that hit industries like coal mining and manufacturing in the US except at a much faster rate and with different driving and resisting forces in automation.

Programmer / engineer compensation in the US market at least has been looking bimodal for at least 12 years now and the first one looks even more devastated in terms of labor than the big tech companies based upon (lack of) job postings and from browsing my connections on LinkedIn.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#638

Earlier quoted context omitted.

>What happened was mathematicians at openAI learned of an imminent development on this problem, and the insight that it entailed, then they were able to prompt a system in the correct direction and spend 20 million dollars to write down the final steps. That's not what happened.

Actually it was. But thanks for elaborating.

No it wasn't. They didn't know 'what direction' to take, and the solution they posted was not in fact the direction the authors took, so evidently you don't know what you're talking about. And they say LLMs hallucinate.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#640

Earlier quoted context omitted.

It requires general intelligence and we don't even have a good understanding of how our's works or a particularly good way of quantifying it. The counter argument is of course maybe you don't need to understand our kind of intelligence to create a different kind and that could well be true but then how do you determine if a system is intelligent. Unless the new system is intelligent enough to reason with us on our le…

Why do you think it requires general intelligence? The parent and I aren't being obtuse here: the history of artificial intelligence research is littered with examples of humans confidently declaring that task X requires general intelligence, then getting humiliated by a neural network doing task X better than humans a few years later. See Go, driving, art (you can complain about the quality of AI art, but it's winni…

I think AI wins against people when we start using KPIs in an industrial view of software like defects / LOC, decomposability, etc), or possibly even maintainability / understandability of a codebase. Gosh, even since the 1980s expert systems were already out-performing entire doctors' boards in diagnosing issues in patients so clearly technical performance in KPIs isn't the only measure by which we adopt technology in society.

When it comes to judgment calls for technical decisions, a lot of interesting innovations appear to come from rejecting conventions / averages in favor of a different set of constraints as a trade-off because we challenge the assumptions we make about the demands being asked of a solution / product. I'm thinking in the constellation of the apocryphal Steve Jobs quote about rejecting asking horse riders what they want because if we asked them they'd ask for a more reliable horse.

Post reply on HN