Live data from Hacker News

Ask HN: Have LLMs Plateaued?

news.ycombinator.com

21–30 of 35 posts

Re: Ask HN: Have LLMs Plateaued?

#21
post #3

Obviously not. And if you look at the things LLMs do poorly there's clearly plenty of room for improvement.

LLM shill argument in a nutshell: "while there is room for improvement in catalytic converters, they are improving practically by the month! Within a decade, internal combustion engines will put out air that is so breathable, it could be used for ventilating a maternity ward."

The exhaust from a hydrogen fuel cell car is drinkable water.

However, experts do not recommend drinking it straight from the tailpipe because the water can pick up dust, dirt, and chemical residues.

(it had to be disclaimed!)

Re: Ask HN: Have LLMs Plateaued?

#22

I'm not answering your actual question, but... even if they have totally plateaued technically, there's still some more improvement left to be had from people learning how best to use them (and when not to).

how can you learn to use a tool that keeps changing unpredictably?

Well, I was assuming that what you learned about driving generation X of a tool would be at least mostly applicable to driving generation X+1 of it. That may be false, at least for some values of X.

Re: Ask HN: Have LLMs Plateaued?

#23
I don't think the models are the bottleneck. Instead, how and where we deploy them remains significantly underrated and underutilized.

Like imagine how crazy that you can get human-like intelligence in a small device, we should be able to do more than a chat interface.

Re: Ask HN: Have LLMs Plateaued?

#24
No, LLMs have not plateaued, and each of the latest releases has been a proof of that.

Fable 5 is a model you instantly FEEL how smart and superior it is. GPT 5.6 Sol is a HUGE incremental improvement in multiple directions and dimensions.

DeepSeek v4 Flash 0731 is a huge improvement over the preview version, using exactly the same architecture.

Kimi K3 gets open weight models very, very close to the frontier.

No my friend, we are not done yet.

Re: Ask HN: Have LLMs Plateaued?

#25
How much more productivity will we be able to squeeze from them?

Probably, depends on how you measure productivity.

If you measure productivity in terms of number of automated bureaucratic events (e.g. creating files, organizing files, generating lines of code, finding bugs in code, generating emails, responding to email) then yes productivity will continue to increase because LLM's are the killer app for increasing bureaucratic events.

If you measure productivity in terms of changes to the material world (e.g. traditional things like trade goods, buildings, food stuffs, irrigation systems, transportation networks, etc.) then no because LLM's have little meaningful impact on those activities...no AGI is going to harvest lettuce for our wedge salads).

Re: Ask HN: Have LLMs Plateaued?

#27
post #24

No, LLMs have not plateaued, and each of the latest releases has been a proof of that. Fable 5 is a model you instantly FEEL how smart and superior it is. GPT 5.6 Sol is a HUGE incremental improvement in multiple directions and dimensions. DeepSeek v4 Flash 0731 is a huge improvement over the preview version, using exactly the same architecture. Kimi K3 gets open weight models very, very close to the frontier. No my…

> you instantly FEEL how smart and superior it is

I felt taken by the change in the system model personality and writing style compared to opus, but I also found it to be much less impressive than I was expecting - let alone that the cost was incredibly high when not given for free.

Are you sure your reaction is not primarily to the improved ergonomics of Fable?

Re: Ask HN: Have LLMs Plateaued?

#29

How much more productivity will we be able to squeeze from them? Probably, depends on how you measure productivity. If you measure productivity in terms of number of automated bureaucratic events (e.g. creating files, organizing files, generating lines of code, finding bugs in code, generating emails, responding to email) then yes productivity will continue to increase because LLM's are the killer app for increasing…

The advances in robotics makes it look like a humanoid robot that could harvest lettuce is only a decade or two away. Making the robot is the easy part. The code to drive it has been the hard part. Until now, that is.

Re: Ask HN: Have LLMs Plateaued?

#30

How much more productivity will we be able to squeeze from them? Probably, depends on how you measure productivity. If you measure productivity in terms of number of automated bureaucratic events (e.g. creating files, organizing files, generating lines of code, finding bugs in code, generating emails, responding to email) then yes productivity will continue to increase because LLM's are the killer app for increasing…

The advances in robotics makes it look like a humanoid robot that could harvest lettuce is only a decade or two away. Making the robot is the easy part. The code to drive it has been the hard part. Until now, that is.

California has about 800,000 agricultural laborers and about 400,000 FTE’s.

That’s a lot of robots…and they probably won’t be able to drive themselves from field to field.

Post reply on HN