Live data from Hacker News

I think you might be fooling yourself with AI

louwrentius.com

131–140 of 174 posts

Re: I think you might be fooling yourself with AI

#131
post #94

Earlier quoted context omitted.

> For things you're already highly skilled at LLMs can be handy but the overall gains, after all is accounted for, are not so clear. Preposterous. I can tell 5.6 Sol to do something like "optimize this entire subsystem to have no ongoing memory allocations, fix these 5 bugs, implement these 2 features" and go work on other code or do something else entirely, and it does it all flawlessly. All things I could have done…

This is going to sound accusatory but I promise it isn’t. Did you feel the same way about 5.5? The same general sentiment keeps being expressed with every model release. The prior generation is immediately cast aside with a vague “well yes, actually we didn’t mean it last time” attitude. I think I even heard Theo the T3 guy suggesting that his work comes to halt if he can’t use the latest generation model, despite he…

Prior to 5.5 I thought the models were very useful but not trustworthy for certain tiers of tasks and needed more in depth verification. Frequent hand edits and times where it felt more productive to just finish something myself than keep rolling the dice on the model not subtly messing it up. But still great for targeted things.

5.5 on Extra High I found to be a huge leap in what it could tackle and how effectively but only up to a point. Still had common failure modes here and there but a big improvement

When I tried 5.6 I wasn't expecting more than an incremental upgrade, but it blew me away with how "smart" it seemed, how it completely sidestepped failure modes I'd been used to dealing with in 5.5, its success rate on the first try for wide-sweeping architectural tasks, and how I've basically never seen a bad hallucination make it through to the final output. It's the first time I've felt like if I left it with a task on the highest setting and it had any way to verify its output, and told it to keep going, it could solve almost anything. This could all be incremental upgrades in many different areas of the model+harness, but the cumulative effect feels like crossing that "escape velocity" threshold. The point where there are no more rough edges and it "just works".

I can't speak to others, but my feelings on all of these have been consistent since I tried them in the first place. I've also heard model providers draw people in at each new release with the model at full performance and then degrade it over time hoping users won't notice, which has seemed plausible anecdotally but I have no hard proof.

Re: I think you might be fooling yourself with AI

#132
post #93

Earlier quoted context omitted.

> For things you're already highly skilled at LLMs can be handy but the overall gains, after all is accounted for, are not so clear. Preposterous. I can tell 5.6 Sol to do something like "optimize this entire subsystem to have no ongoing memory allocations, fix these 5 bugs, implement these 2 features" and go work on other code or do something else entirely, and it does it all flawlessly. All things I could have done…

> it does it all flawlessly Serious question: How do you assess this? It seems to me that it would take very significant time to establish that conclusion.

Validation is really hard and requires a great understanding of the entire relevant code. So to answer your question, they don’t know or asked the AI to do that too which has very obvious problems.

I find it hard to believe they would spend the effort and money making an AI do tasks like that, but are ALSO willing to spend the effort and time to properly validate the changes. I’m sure many do… but probably not a large percentage.

Re: I think you might be fooling yourself with AI

#133
post #34

> Meanwhile, a study (late 2025) seems to report that although participating developers felt they completed tasks faster using AI, they where around 19% slower. METR reran the study early this year and, while they caveat it, this time they found a speedup, which is consistent with subjective estimates of productivity also having increased -- the simplest explanation is that subjective estimates exaggerate, but there'…

Have you looked at the raw data of the follow-up? It doesn't paint a pretty picture, many tasks attempted but not completed, refusal to attempt without AI, higher overall estimates and completion times than the previous study (both for AI and Non AI).

There's no definitive conclusion about productivity that can be drawn from the followup. However I think it's reasonable to argue that AI use overtime has made devs worse without AI.

Re: I think you might be fooling yourself with AI

#134
post #93

Earlier quoted context omitted.

> it does it all flawlessly Serious question: How do you assess this? It seems to me that it would take very significant time to establish that conclusion.

Validation is really hard and requires a great understanding of the entire relevant code. So to answer your question, they don’t know or asked the AI to do that too which has very obvious problems. I find it hard to believe they would spend the effort and money making an AI do tasks like that, but are ALSO willing to spend the effort and time to properly validate the changes. I’m sure many do… but probably not a larg…

In my case it's usually a video game in question, not a life support system, so I'm not poring over every line of code.

I mainly look at:

- The overall architecture to make sure it's in line with what I want

- The points where it interacts with certain other systems I'm concerned about

And I test the feature myself in addition to the test cases. Sol has proven to be outstanding at good test cases and having a coherent view on overall architecture integration. I basically never think to myself "this entire implementation is slop, I need to rip it out and refactor it" anymore.

I suspect the model's capacity to translate requirements into contextually appropriate test cases, and insistence on thoroughly covering functionality in tests, is a large part of its strength.

Re: I think you might be fooling yourself with AI

#135
post #9

These articles are so off from reality to the point that it can't be taken seriously at all. I keep wondering how they keep popping at HN. At least the author admit he is biased..

I think there's a simple dichotomy based on use cases. For things you're already highly skilled at LLMs can be handy but the overall gains, after all is accounted for, are not so clear. But for things you aren't good at, they're zomg amazing. So for somebody evaluating things based on what they do at work (where they're probably quite competent) or on personal projects well within their own domain, then it's 'hey wha…

That's one dimension. Another is how much you care.

For my own projects, little productivity apps for my own needs, I'm easily 20x more productive. I'm building so many more of them, and at a much higher level of polish. But that's simply because I don't care about the code, long-term maintenance, or anything else. I just need them to do a thing, and if they do the thing, I'm good. I used to make an effort calculation before building these things, but now I just tell Claude "hey, build this", and two hours later it's built well enough to be useful.

OTOH, at work, I'm probably about as fast as I was before, maybe even a bit slower. But the quality of my work has increased quite a bit, because I no longer take the shortcuts I used to take to push things out. I'm spending much more time in planning and in code reviews rather than actually typing code.

Either way, the idea that these obvious changes are just people fooling themselves is, at this point, no longer a reasonable position. And the claim that "AI will run out of money and just be turned off due to the huge operational cost" is so implausible that I would feel ashamed of myself if I used this as a straw man for what AI skeptics believe.

Re: I think you might be fooling yourself with AI

#136

These AI sceptics always omit the possibility that running LLMs could become much cheaper in the future, which it almost certainly will. This is largely an infrastructure problem, and it will be solved just as similar problems were solved with storage, network capacity, and computing power. There will be bankruptcies, but the technology will prevail, just as the web did after the dot com bubble. Personally, I see the…

This mostly reads like wishful thinking about the future, that's not a basis to make decisions on right now.

I do not think that wishful thinking is assumption that rapidly evolving hardware which is currently one of the bottlenecks in high demand would become faster in the future.

Re: I think you might be fooling yourself with AI

#137

Earlier quoted context omitted.

Are you claiming you’ve never been responsible for a bug that caused data corruption? Perhaps not, but it’s hardly a novel post-LLM bug. It’s regular stuff and obviously you know that.

As a 2 man developer team working on a mono repo with clear and simple architecture, no I have never caused urgent data corruption or seen it happen. That is more something that happens in codebases that are too big and complex to fully grasp. That it has already happened to you in this scenario does not bode well for the future stability of your codebase, in my opinion.

Urgent is an exaggeration I suppose. I have a high standard of customer service so to me it’s urgent, but the data was corrupted for 3 months before anyone noticed so it wasn’t critical.

Ok I guess we just suck, whatever. We and our customers are very happy.

Re: I think you might be fooling yourself with AI

#138

Earlier quoted context omitted.

I think there's a simple dichotomy based on use cases. For things you're already highly skilled at LLMs can be handy but the overall gains, after all is accounted for, are not so clear. But for things you aren't good at, they're zomg amazing. So for somebody evaluating things based on what they do at work (where they're probably quite competent) or on personal projects well within their own domain, then it's 'hey wha…

> For things you're already highly skilled at LLMs can be handy but the overall gains, after all is accounted for, are not so clear. Preposterous. I can tell 5.6 Sol to do something like "optimize this entire subsystem to have no ongoing memory allocations, fix these 5 bugs, implement these 2 features" and go work on other code or do something else entirely, and it does it all flawlessly. All things I could have done…

That's highly unlikely that this is done flawlessly. This morning I found a bug in our error handling, I asked Sol to fix it. He rewrote the way the specific component that failed called the exception. I asked it to retry and go deeper, inside the library. It solution was to ignore the part of the message that caused the issue. Which worked very well, but wasn't the root cause (it was a serialization issue. It's always a serialization issue). It couldn't go deeper, always stuck trying to catch the error rather than fixing it for real. We found the root cause, fixed it, and it was done, one line change and 15 lines of test, in one hours. Instead of four hours of nothing except spending tokens.

But truly I'm mad at the vscode harness. I used to be able to follow the 'thought' of the AI model and steer them when I saw a mistake, it's now way harder to do as multiple research tasks are done in parallel.

Re: I think you might be fooling yourself with AI

#139
post #138

Earlier quoted context omitted.

> For things you're already highly skilled at LLMs can be handy but the overall gains, after all is accounted for, are not so clear. Preposterous. I can tell 5.6 Sol to do something like "optimize this entire subsystem to have no ongoing memory allocations, fix these 5 bugs, implement these 2 features" and go work on other code or do something else entirely, and it does it all flawlessly. All things I could have done…

That's highly unlikely that this is done flawlessly. This morning I found a bug in our error handling, I asked Sol to fix it. He rewrote the way the specific component that failed called the exception. I asked it to retry and go deeper, inside the library. It solution was to ignore the part of the message that caused the issue. Which worked very well, but wasn't the root cause (it was a serialization issue. It's alwa…

Somehow I've never had an experience like that on Sol. I find it to never really "cheat" like that. I've had all my best results from Codex CLI though, maybe that's somehow different. I also don't go below xhigh for anything serious and if it's having an issue I escalate the thinking mode.

Re: I think you might be fooling yourself with AI

#140

Earlier quoted context omitted.

Dig down and I think you’ll find that AI boosters are really quite insecure.

It certainly seems that way from this - admittedly tiny - sample

As an anti-AI booster, I admit that I'm quite insecure, too. And angry. But man, these endless threads with folks endlessly hyping their productivity gains and haranguing others for going against the grain are just blatantly pathological. I feel like I'm an in an episode of Pluribus.
Post reply on HN