Earlier quoted context omitted.
I agree -- skepticism is totally healthy. And there are so many great ways to poke holes in the true underlying narratives (not the headlines that people seem to pull from). E.g. evaluation science is a wasteland (not for wont of very smart people trying very hard to get them right). How do we tackle the power requirements in a way that is sustainable? Etc. etc. But stuff like this im not sure I understand: > It's a…
>> if its a great tool, then how is it _only_ being used to "feed the greed" and what do you mean by that? Look around?
2025: The Year in LLMs
551–560 of 643 posts
Re: 2025: The Year in LLMs
#552Earlier quoted context omitted.
Seems like Nvidia will be focusing on the super beefy GPUs and leaving the consumer market to a smaller player
AMD owns a lot of the consumer market already; handhelds, consoles, desktop rigs and mobile ... they are not a small player.
Re: 2025: The Year in LLMs
#553Earlier quoted context omitted.
Yeah the internet kind of started with ARPANET in 1969 and didn't really get going with the public till around 1999 so thirty years on. Here's a graph of internet takeoff with Krugman's famous quote of 1998 that it wouldn't amount to much being maybe the end of the skepticism https://www.contextualize.ai/mpereira/paul-krugmans-poor-pre... In common with AI there was probably a long period when the hardware wasn't rea…
Thats all irrelevant. Is/was there tremendous value to be had by being able to transport data? Of course. No doubt about it. Everything else got figured out and investments were made because of that. The same line of thinking does not hold with LLMs given their non-deterministic nature. Time will tell where things land.
Re: 2025: The Year in LLMs
#554Earlier quoted context omitted.
2022/2023: "It hallucinates, it's a toy, it's useless." 2024/2025: "Okay, it works, but it produces security vulnerabilities and makes junior devs lazy." 2026 (Current): "It is literally the same thing as a psychic scam." Can we at least make predictions for 2027? What shall the cope be then! Lemme go ask my psychic.
I suppose it's appropriate that you hallucinated an argument I did not make, attacked the straw man, and declared victory.
I have no idea if those two points, ML and brains, are just different points on the same Pareto frontier of some useful metrics, but I am increasingly suspecting they might be.
Re: 2025: The Year in LLMs
#555Earlier quoted context omitted.
2022/2023: "It hallucinates, it's a toy, it's useless." 2024/2025: "Okay, it works, but it produces security vulnerabilities and makes junior devs lazy." 2026 (Current): "It is literally the same thing as a psychic scam." Can we at least make predictions for 2027? What shall the cope be then! Lemme go ask my psychic.
2022/2023: "Next year software engineering is dead" 2024: "Now this time for real, software engineering is dead in 6 months, AI CEO said so" 2025: "I know a guy who knows a guy who built a startup with an LLM in 3 hours, software engineering is dead next year!" What will be the cope for you this year?
… to one of the models in Jan 2024 being able to repeatedly add features to the same single-page web app without corrupting its own work or hallucinating the APIs it had itself previously generated…
… to last month using a gifted free week of Claude Code to finish one project and then also have enough tokens left over to start another fresh project which, on that free left-over credit, reached a state that, while definitely not well engineered, was still better than some of the human-made pre-GenAI nonsense I've had to work with.
Wasn't 3 hours, and I won't be working on that thing more this month either because I am going to be doing intensive German language study with the goal of getting the language certificate I need for dual citizenship, but from the speed of work? 3 weeks to make a startup is already plausible.
I won't say that "software engineering" is dead. In a lot of cases however "writing code" is dead, and the job of the engineer should now be to do code review and to know what refactors to ask for.
Re: 2025: The Year in LLMs
#556Earlier quoted context omitted.
2022/2023: "Next year software engineering is dead" 2024: "Now this time for real, software engineering is dead in 6 months, AI CEO said so" 2025: "I know a guy who knows a guy who built a startup with an LLM in 3 hours, software engineering is dead next year!" What will be the cope for you this year?
The cope + disappointment will be knowing that a large population of HN users will paint a weird alternative reality. There are a multitude of messages about AI that are out there, some are highly detached from reality (on the optimistic and pessimistic side). And then there is the rational middle, professionals who see the obvious value of coding agents in their workflow and use them extensively (or figure out how t…
This is a surprising claim. There's only 3 orders of magnitude between US data centre electricity consumption and worldwide primary energy (as in, not just electricity) production. Worldwide electricity supply is about 3/20ths of world primary energy, so without very rapid increases in electricity supply there's really only a little more than 2 orders of magnitude growth possible in compute.
Renewables are growing fast, but "fast" means "will approach 100% of current electricity demand by about 2032". Which trend is faster, growth of renewable electricity or growth of compute? Trick question, compute is always constrained by electricity supply, and renewable electricity is growing faster than anything else can right now.
Re: 2025: The Year in LLMs
#557Earlier quoted context omitted.
Sure: at my MAANG company, where I watch the data closely on adoption of CC and other internal coding agent tools, most (significant) LOC are written by agents, and most employees have adopted coding agents as WAU, and the adoption rate is positively correlated with seniority. Like a lot of things LLM related (Simon Willison's pelican test, researchers + product leaders implementing AI features) I also heavily "vibe"…
We probably work at the same company, given you used MAANG instead of FAANG. As one of the WAU (really DAU) you’re talking about, I want to call out a couple things: 1) the LOC metrics are flawed, and anyone using the agents knows this - eg, ask CC to rewrite the 1 commit you wrote into 5 different commits, now you have 5 100% AI-written commits; 2) total speed up across the entire dev lifecycle is far below 10x, mos…
If the M stands for Meta, I would also like to note that as a user, I have been seeing increasingly poor UI, of the sort I'd expect from people committing code that wasn't properly checked before going live, as I would expect from vibe coding in the original sense of "blindly accept without review". Like, some posts have two copies of the sender's name in the same location on screen with slightly different fonts going out of sync with each other.
I can easily believe the metrics that all [MF]AANG bonuses are denominated in are going up, our profession has had jokes about engineers gaming those metrics even back when our comics were still printed in books: https://imgur.com/bug-free-programs-dilbert-classic-tyXXh1d
Re: 2025: The Year in LLMs
#558Earlier quoted context omitted.
To be fair, I’ll take a non-biased 16 person study over “internal measures” from a MAANG company that burned 100s of billions on AI with no ROI that is now forcing its employees to use AI.
What do you think about the METR 50% task length results? About benchmark progress generally?
The converse of this is that if those tasks are representative of software engineering as a whole, I would expect a lot of other tasks where it absolutely sucks.
This expectation is further supported by the number of times people pop up in conversations like this to say for any given LLM that it falls flat on its face even for something the poster thinks is simple, that it cost more time than it saved.
As with supposedly "full" self driving on Teslas, the anecdotes about the failure modes are much more interesting than the success: one person whose commute/coding problem happens to be easy, may mistake their own circumstances for normal. Until it does work everywhere, it doesn't work everywhere.
When I experiment with vibe coding (as in, properly unsupervised), it can break down large tasks into small ones and churn through each sub-task well enough, such that it can do a task I'd expect to take most of a sprint by itself. Now, that said, I will also say it seems to do these things a level of "that'll do" not "amazing!", but it does do them.
But I am very much aware this is like all the people posting "well my Tesla commute doesn't need any interventions!" in response to all the people pointing out how it's been a decade since Musk said "I think that within two years, you'll be able to summon your car from across the country. It will meet you wherever your phone is … and it will just automatically charge itself along the entire journey."
It works on my [use case], but we can't always ship my [use case].
Re: 2025: The Year in LLMs
#559Earlier quoted context omitted.
> What's my motivation? Are you really going to insult my and others' intelligence like this? Directly or indirectly, your motivation is money. You already offer monthly subscriptions to your blog, and you're clearly trying to build a monetizable brand for yourself as a leading authority on AI, especially as it pertains to software development.
If my motivation was money I would cash in on the reputation I've already built and go and land a Silicon Valley salaried job somewhere. Sponsorship from my monthly newsletter doesn't come close. Seriously, do you have any idea how much money I'm leaving on the table right now NOT having a real job in this space? Being a blogger is wildly financially irresponsible!
Re: 2025: The Year in LLMs
#560Earlier quoted context omitted.
Honestly, I wouldn't be surprised if a system that's an LLM at its core can attain AGI. With nothing but incremental advances in architecture, scaffolding, training and raw scale. Mostly the training. I put less and less weight on "LLMs are fundamentally flawed" and more and more of it on "you're training them wrong". Too many "fundamental limitations" of LLMs are ones you can move the needle on with better training…
That depends on how you define AGI - it's a meaningless term to use since everyone uses it to mean different things. What exactly do you mean ?! Yes, there is a lot that can be improved via different training, but at what point is it no longer a language model (i.e. something that auto-regressively predicts language continuations)? I like to use an analogy to the children's "Stone Soup" story whereby a "stone soup" (…
You described modern RLVR for tasks like coding. Plug an LLM into a virtual env with a task. Drill it based on task completion. Force it to get better at problem-solving.
It's still an autoregressive next token prediction engine. 100% LLM, zero architectural changes. We just moved it past pure imitation learning and towards something else.