Earlier quoted context omitted.
Framed this way, it's useful for saving time creating or finding those snippets at least.
for …. Yeh, you are right. A snippet for they would have taken longer than booting an agent and let it hallucinate for—loop parameters that didn’t even exist.
Where I'm at with AI
61–70 of 72 posts
Re: Where I'm at with AI
#62Earlier quoted context omitted.
First car prototypes were useless and it has taken a few decades to have a good version. The first combustion engine was in 1826. Would you buy a prototype or a carriage for transportation at that time?
no, but AI isn't going to light on fire as I drive and potentially kill me. it's also not an exorbitant expense.
Re: Where I'm at with AI
#63I'm very PRO AI and a daily user but let's be real, where is the productivity gain? It's at the edges, it's in greenfield, we've seen a handful of "things written with AI" be launched successfully and the majority of the time these are shovels for other AI developers. Large corporations are not releasing even 5x higher velocity or quality. Ditto for small corporations. If the claimed multipliers were broadly real at…
This has been the big one for me. Of all the best devs I know, only one is really big on AI, but he’s also the one that spent years in a relational database without knowing what isolation levels are. He was absolutely a “greenfield” kinda dev - which he was great at, tons of ideas but not big on the small stuff.
The rest use them judiciously and see a lot of value, but reading and understanding code is still the bottleneck. And reviewing massive PRs is not efficient.
Re: Where I'm at with AI
#64This article is different because it actually talks about code review, which I don’t see very often. Especially in the ultra-hype 1000x “we built an operating system in a day using agentic coding”, it seems like code review is implied to be a thing of the past. As long as code continues to need to be reviewed to (mostly) maintain a chain of liability, I don’t see SWE going anywhere like both hypebros and doomers seem…
1. You know exactly what the code should look like ahead of time. Agents are handy here, but this is rare and takes no time at all as is
2. You don’t know exactly what it should be. Building the mental model by hand is faster than working backwards from finished work. Agents help with speeding up your knowledge here, but having them build the model is not good
Code review is and always was the bottle neck. You are doing code review when you write code.
Re: Where I'm at with AI
#65Earlier quoted context omitted.
First car prototypes were useless and it has taken a few decades to have a good version. The first combustion engine was in 1826. Would you buy a prototype or a carriage for transportation at that time?
If you couldn't foresee how they would eventually be useful with improvements over time you probably bought a lot of horse carriages in 1893 and appropriately lost your ass.
How to avoid being a Duryea, a Knox, a Marsh, a Maxwell-Briscoe, or a Pope-Toledo seems to be the real question.
Same thing pretty much happened in the early days of radio, with the addition of vicious patent wars. Which I'm sure we'll eventually see in the AI field, once the infinite money hose starts to dry up.
Re: Where I'm at with AI
#66Earlier quoted context omitted.
Yea I hear this a lot, do people genuinely dismiss that there has been step change progress over 6-12 months timescale? I mean it’s night and day, look at benchmark numbers… “yea I don’t buy it” ok but then don’t pretend you’re objective
I think I'd be in the "don't buy it" camp, so maybe I can explain my thinking at least. I don't deny that there's been huge improvements in LLMs over the last 6-12 months at all. I'm skeptical that the last 6 months have suddenly presented a 'category shift' in terms of the problems LLMs can solve (I'm happy to be proved wrong!). It seems to me like LLMs are better at solving the same problems that they could solve 6…
Yes I agree here in principle here in some cases: I think there are certainly problems that LLMs are now better at but that don't reach the critical reliability threshold to say "it can do this". E.g. hallucinations, handling long context well (still best practice to reset context window frequently), long-running tasks etc.
> That's kind of a fuzzier point, and a hard one to know until we all have hindsight. But I think OP is right that people have been claiming "LLMs are fundamentally in a different category to where they were 6 months ago" for the last 2 years - and as yet, none of those big improvements have yet unlocked a whole new category of use cases for LLMs.
This is where I disagree (but again you are absolutely right for certain classes of capabilities and problems).
- Claude code did not exist until 2025
- We have gone from e.g. people using coding agents for like ~10% of their workflow to like 90-100% pretty typically. Like code completion --> a reasonably good SWE (with caveats and pain points I know all too well). This is a big step change in what you can actually do, it's not like we're still doing only code completion and it's marginally better.
- Long horizon task success rate has now gotten good enough that basically also enable the above (good SWE) for like refactors, complicated debugging with competing hypotheses, etc, looping attempts until success
- We have nascent UI agents now, they are fragile but will see a similar path as coding which opens up yet another universe of things you can only do with a UI
- Enterprise voice agents (for like frontline support) now have a low enough bounce rate that you can actually deploy them
So we've gone from "this looks promising" to production deployment and very serious usage. This may kind of be like you say "same capabilities but just getting gradually better" but at some point that becomes a step change. Before a certain failure rate (which may be hard to pin down explicitly) it's not tolerable to deploy, but as evidenced by e.g. adoption alone we've crossed that threshold, especially for coding agents. Even sonnet 4 -> opus 4.5 has for me personally (beyond just benchmark numbers) made full project loops possible in a way that sonnet 4 would have convinced you it could and then wasted like 2 whole days of your time banging your head against the wall. Same is true for opus 4.5 but its for much larger tasks.
> To be honest, it's a very tricky thing to weight into, because the claims being made around LLMs are very varied from "we're 2 months away from all disease being solved" to "LLMs are basically just a bit better than old school Markov chains". I'd argue that clearly neither of those are true, but it's hard to orient stuff when both those sides are being claimed at the same time.
Precisely. Lots and lots of hyperbole, some with varying degrees of underlying truth. But I would say: the true underlying reality here is somewhat easy to follow along with hard numbers if you look hard enough. Epoch.ai is one of my favorite sources for industry analysis, and e.g. Dwarkesh Patel is a true gift to the industry. Benchmarks are really quite terrible and shaky, so I don't necessarily fault people "checking the vibes", e.g. like Simon Willison's pelican task is exactly the sort of thing that's both fun and also important!
Re: Where I'm at with AI
#67Earlier quoted context omitted.
Well considering people that disagree with you “shills” is maybe a bad start and indicates you kind of just have an axe to grind. You’re right that there can be serious local issues for data centers but there are plenty of instances where it’s a clear net positive. There’s a lot of nuance that you’re just breezing over and then characterizing people that point this out as “shills”. Water and electricity demands do no…
It’s already causing massive problems before we even got to models that even do anything actually useful. Wait until other countries jump in the bandwagon and see energy prices jump. Currently it is mostly the US and China. And RAM and GPU prices are already through the roof with SSDs to follow. That is for now with ONLY the US market mostly. And those “net benefits” you talk about have very questionable data behind…
RAM and GPU prices going up, sure ok, but again: if you're claiming there is no net benefit from AI what is your evidence for that? These contracts are going through legally, so what basis do you have to prevent them from happening? Again I say its site specific: plenty of instances where people have successfully prevented data centers in their area, and lots of problems come up (especially because companies are secretive about details so people may not be able to make informed judgements).
What are the massive problems besides RAM + GPU prices, which again, what is the societal impact from this?
Re: Where I'm at with AI
#68This article is different because it actually talks about code review, which I don’t see very often. Especially in the ultra-hype 1000x “we built an operating system in a day using agentic coding”, it seems like code review is implied to be a thing of the past. As long as code continues to need to be reviewed to (mostly) maintain a chain of liability, I don’t see SWE going anywhere like both hypebros and doomers seem…
This is one of the key elements that will shift the SWE discipline. We will move to more explicitly editing systems by their footprint and guarantees, and verifying that what exists matches those.
Shareholders don’t want a “release” button that when they press it there’s a XX% chance that the entire software stack collapses because an increasingly complex piece of software is understood by no one but a black box called “AI” that can’t be fired or otherwise held accountable.
Re: Where I'm at with AI
#69> increases in worker productivity at best increase demand for labor, and at worst result in massive disruption - they never result in the same pay for less manual work. Exactly. Strongly agree with that. This closed world assumption never holds. We would only do less work if nothing else changes. But of course everything changes when you lower the price of creating software. It gets a lot cheaper. So, now you get a…
> There's a reason this stuff was never upgraded: doing so is super expensive. Now that just got a bit cheaper to do. Maybe we'll get around to doing some of that. I don’t think typing code was ever the problem here. I’m doubtful this got any cheaper.
Re: Where I'm at with AI
#70Earlier quoted context omitted.
If you couldn't foresee how they would eventually be useful with improvements over time you probably bought a lot of horse carriages in 1893 and appropriately lost your ass.
The problem is, that's only one way you could have lost your ass during the transition from horses to cars. I think that's what skydhash was getting at. There were hundreds of early car companies, all competing aggressively with one another, each with something to sell that the others seemed to lack. For every one that succeeded, dozens ended in bankruptcy, and not necessarily through any obvious fault of their own.…