Live data from Hacker News

Ask HN: Why is it taken for granted that LLM models will keep improving?

news.ycombinator.com

41–50 of 68 posts

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#41
I'm curious if limits of like, thermodynamics, won't play a part here. Or maybe also ecological limits: how long will we allow corporations to use essential, scarce resources to train models without paying their fair share? [0]

I'm not an expert here either but I wonder if there will be the same "leap" we saw from ChatGPT3-4 or if there's a diminishing curve to performance, ie: adding another trillion parameters has less of a noticeable effect than the first few hundred billion.

[0] https://fortune.com/2023/09/09/ai-chatgpt-usage-fuels-spike-... -- I am fairly certain they paid for that water, it was not a commensurate price given the circumstances, and if they had to ask to use it first the answer would have been, no, by a reasonable environmental stewardship organization.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#42
post #6

Because a lot of smart people are spending a lot of time, money, and effort on this. It's as simple as that. We could go into all sorts of details, like how increase in GPU capabilities will improve training capabilities, both in size and speed, or how GPU(/TPU) capabilities will improve, or how better techniques will make training on the same data set result in better models, or where other improvements will make be…

I mean a lot of highly payed/intelligent people worked in crypto/fusion/quantum computing, still all of those topics are evolving rather gradually.

We know why those are.

Cryptocurrency’s need to be fully decentralised is the thorn in it’s side. Be your own bank is a bit too much for most people used to cash or a bank account they can call up if there is a problem. It has fundamental social problems that there may be solutions to but probably not.

Fusion and quantum are massive physics and engineering challenges. With ML we are already building the chips to scale, so we know it is scalable and doable.

It is a 50 to 100 problem not 0 to 1.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#43
What's not mentioned here is test-time compute. Idea being that, sure, you can spend a ton of compute power on pre-training and fine-tuning, but generation is difficult. So instead of spending all time and power more focused on that, how about spending some time and power on it for the model to generate a bunch of possibilities, and then spend the rest of time having a model verify what's been generated for correctness. That's the Let's Verify Step by step.

Great video to talk about this: https://www.youtube.com/watch?v=ARf0WyFau0A

In threads on LLMs, this point doesn't get brought up as much as I'd expect, so I'm curious if I'm missing talks on this or maybe it's wrong. But I see this as the way forward. Models generating tons of answers, and other models being able to pick out the correct ones, and the combinations being beyond human ability, where after, humans can do their own verification.

Edit:

Think of it this way. Trying to create something isn't easy. If I was to write a short story, it'd be very difficult, even if I spent years reading what others have written to learn their patterns. If I then tried to write and publish a single one myself, no chance it'd be any good.

But _judging_ short stories is much easier to do. So if I said screw it, I'll read a couple stories to get the initial framework, then write 100 stories in the same amount of time I'd have spent reading and learning more about short stories, I can then go through the 100 and pick out the one I think is the best and publish that.

That's where I see LLMs going and what the video and papers mentioned in the video say.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#44

I dunno if LLMs will get better, but ML in general is a task of compression, and there is definitely a whole bunch of human knowledge and history that neural nets can compress. Its not unfeasable in the future to have a box at home that you can ask a fairly complicated question, like "how do I build a flying car", and it will have the ability to - tell you step by step instructions of what you need to order - write a…

unbounded possibilities! imagine, and hear me out, asking the box: 'how do I build a box that can answer fairly complicated questions', and getting the output in an automatized way all the way from atoms, ready to be plugged into the power grid.

It's a wishing for more wishes from a genie scenario. But it isn't enough to just have a box that can do this, the box needs a body, preferably a humanoid one so that it can interface easily with our existing infrastructure.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#45
I don't think it's a universal assumption. Some people do think it will hit a wall (and maybe do so soon), others think it can keep improving easily by scaling up the compute or the training data.

Good LLMs like ChatGPT are a relatively new technology so I think it's hard to say either way. There might be big unrealized gains by just adding more compute, or adding/improving training data. There might be other gains in implementation, like some kind of self-improvement training, a better training algorithm, a different kind of neural net, etc. I think it's not unreasonable to believe there are unrealized improvements given the newness of the technology.

On the other hand, there might be limitations to the approach. We might never be able to solve for frequent hallucinations, and we might not find much more good training data as things get polluted by LLM output. Data could even end up being further restricted by new laws meaning this is about the best version we will have and future versions will have worse input data. LLMs might not have as many "emergent" behaviors as we thought and may be more reliant on past training data than previously understood, meaning they struggle to synthesize new ideas (but do well at existing problems they've trained on). I think it's also not unreasonable to believe LLMs can't just improve infinitely to AGI without more significant developments.

Speculation is always just speculation, not a guarantee. We can sometimes extrapolate from what we've seen, but sometimes we haven't seen enough to know the long term trend.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#46

The scaling laws (the original Kaplan paper, Chinchilla, and OpenAI's very opaque scaling graphs for GPT-4) suggest indefinite improvement for the current style of transformers with additional pre-training data and parameters. No one has hit a model/dataset size where the curves break down, and they're fairly smooth. Usually simple models that accurately predict performance work pretty well nearby existing performanc…

> No one has hit a model/dataset size where the curves break down, and they're fairly smooth.

The same was true of transistors, until it wasn't and they started diverging from the predictions about how they would behave when very small. Sometime around the late Netburst era (the Pentium 4/Netburst architecture was sunk by this problem - they assumed, designing it, that it would scale to 8-10GHz on a sane power budget, and it simply didn't as the "improvement per transistor shrink" became less and less).

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#47
post #3

The main assumption of techno-optimism is that a large enough computer can do anything people can do and it can do it better. The goal of techno-optimism is to create a mechanical god that will rule the planet and scaling LLMs is a stepping stone to that goal. I, of course, already know how to do all this for a mere $80B.

"create a mechanical god that will rule the planet" -- on what basis people call this 'optimism'?!

They don't. They say things like "infinite resources for everyone" and "compute/electricity too cheap to meter" and so on and so forth. I have distilled the techno-optimist manifesto down to what it actually looks like in reality, i.e. a global panopticon that controls everything with algorithms.

Per usual, I can build this technological panopticon/utopia for a bargain price of $80B. Some people think it can be done for cheaper but they haven't spent as much time as I have on this problem. I have the architecture ready to go, all I need is the GPUs, cameras, microphones, speakers, and wireless data network. The software is the easy part but the panoptic infrastructure is what requires the most capital. The software/brain can be done for maybe $2B but it needs eyes, ears, and a mouth to actually be useful.

The second stage is building up the actuators to bypass people but once the panopticon is ready it won't be hard to build up the robot factories to enact the will of AGI directly via robots acting on the environment.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#48

What's not mentioned here is test-time compute. Idea being that, sure, you can spend a ton of compute power on pre-training and fine-tuning, but generation is difficult. So instead of spending all time and power more focused on that, how about spending some time and power on it for the model to generate a bunch of possibilities, and then spend the rest of time having a model verify what's been generated for correctne…

Isn't that just GAN, but with LLMs?

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#49
The recent history of bigger LLMs suddenly being capable of new things is kind of miraculous. This blog post is a decent overview: https://blog.research.google/2022/11/characterizing-emergent... " In many cases, the performance of a large language model can be predicted by extrapolating the performance trend of smaller models."

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#50

Earlier quoted context omitted.

I mean a lot of highly payed/intelligent people worked in crypto/fusion/quantum computing, still all of those topics are evolving rather gradually.

We know why those are. Cryptocurrency’s need to be fully decentralised is the thorn in it’s side. Be your own bank is a bit too much for most people used to cash or a bank account they can call up if there is a problem. It has fundamental social problems that there may be solutions to but probably not. Fusion and quantum are massive physics and engineering challenges. With ML we are already building the chips to scal…

and we might know what problems those are for AI in a year or two, look back how crypto was viewed on by some...

I also was not arguing for ai to not improve by quit a bit more.. I actually think ai will make a few more big steps forward, but the garantee for this is not ankered in the inflowing capital/talent but instead in the relative clear path forward of the technologie and partly known ineffective architecture.

Post reply on HN