Live data from Hacker News

Ask HN: Why is it taken for granted that LLM models will keep improving?

news.ycombinator.com

21–30 of 68 posts

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#21
Like others in this thread have said, we're just starting to explore the technology. I view it as akin to early CPUs like the 6502 which only did the absolute minimum to today's monsters with large memory caches, predictive logic, dedicated circuits, thousands of binary calculation shortcuts and more all built in. Each small improvement adds up.

From a software perspective, I've wondered for a while if as LLM usage matures, there will be an effort to optimize hotspots like what happened with VMs, or auto indexing like in relational DBs. I'm sure there are common data paths which get more usage, which could somehow be prioritized, either through pre-processing or dynamically, helping speed up inference.

Also, GPT4 seems to include multiple LLMs working in concert. There's bound to be way more fruit to picked along that route as well. In short, there's tons of areas where improvements large and small can be made.

As always in computer science, the maxim, "Make it work, make it work well, then make it work fast," applies here as well. We're collectively still at step one.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#22

The scaling laws (the original Kaplan paper, Chinchilla, and OpenAI's very opaque scaling graphs for GPT-4) suggest indefinite improvement for the current style of transformers with additional pre-training data and parameters. No one has hit a model/dataset size where the curves break down, and they're fairly smooth. Usually simple models that accurately predict performance work pretty well nearby existing performanc…

This is the most accurate answer so far re. The scaling laws. It has been demonstrated that LLMs follow quite clear power laws with respect to performance. In fact, the performance of any model can be determined from the number of parameters it has and the amount of data it is given. The Wikipedia article on Neural Scaling laws provides a brief, accessible, summary of this. Both data and parameters are expected to increase in coming years, so models are expected to improve.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#23
post #14
post #7

Earlier quoted context omitted.

I did the numbers already using a small LLM and it said the real number is $80B and no ponies.

Would that be $80B in 2020 or 2024 dollars? What vintage is your small LLM?

It accounted for inflation, Trumponomics, Bidenomics, and regular economics. I trust the number, it feels right to me.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#24
post #15

LLMs are comprised of just three elements Data Compute Algorithms All three are just scratching the surface of what is possible. Data: What has been scraped off the internet is just Compute: While 3nm processes are approaching an atomic limit (0.21nm for Si), there is still room to explore more densely packed transistors or other materials like Gallium Nitride or optical computing. Not only that but there is a lot of…

> LLMs are comprised of just three elements

> Data

> Compute

> Algorithms

Not to be facetious but so is all other software. LLMs appear to scale in correlation to the first two but it's not clear what the correlation is and that's the basis of the question being asked.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#25
post #3

The main assumption of techno-optimism is that a large enough computer can do anything people can do and it can do it better. The goal of techno-optimism is to create a mechanical god that will rule the planet and scaling LLMs is a stepping stone to that goal. I, of course, already know how to do all this for a mere $80B.

I think I can do it far, far cheaper than that... and I've been talking about it, in public, for over a decade.[1] I could easily be wrong, of course. It really all depends on how much power a 4x4 LUT and 4 bit latch Latch Leak, and how much energy it takes to clock data through them for a cycle, and how fast they can be cycled. If the number are good, this thing will be amazingly cool. I can't find good numbers anyw…

I think $80B is very cheap because the outcome is going to be a technological utopia. Honestly, I think my price is a bargain deal for the inhabitants of Earth.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#26
post #3

The main assumption of techno-optimism is that a large enough computer can do anything people can do and it can do it better. The goal of techno-optimism is to create a mechanical god that will rule the planet and scaling LLMs is a stepping stone to that goal. I, of course, already know how to do all this for a mere $80B.

I can do it for $20B and a pony

Only because you're cheating by using the neurons of the pony.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#27
post #15

LLMs are comprised of just three elements Data Compute Algorithms All three are just scratching the surface of what is possible. Data: What has been scraped off the internet is just Compute: While 3nm processes are approaching an atomic limit (0.21nm for Si), there is still room to explore more densely packed transistors or other materials like Gallium Nitride or optical computing. Not only that but there is a lot of…

For data though, as LLM's generate more output, over time wouldn't they be expected to mess themselves with their own generated data?

Wouldn't that be the wall we'll hit? Think of how shitted up Google Search is with generated garbage, I'm imagining we're already in the 'golden age' where we were able to train on good datasets before it gets 'polluted' with LLM generated data that may not be accurate, and it just continues to become less accurate over time.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#28
post #15

LLMs are comprised of just three elements Data Compute Algorithms All three are just scratching the surface of what is possible. Data: What has been scraped off the internet is just Compute: While 3nm processes are approaching an atomic limit (0.21nm for Si), there is still room to explore more densely packed transistors or other materials like Gallium Nitride or optical computing. Not only that but there is a lot of…

For data though, as LLM's generate more output, over time wouldn't they be expected to mess themselves with their own generated data? Wouldn't that be the wall we'll hit? Think of how shitted up Google Search is with generated garbage, I'm imagining we're already in the 'golden age' where we were able to train on good datasets before it gets 'polluted' with LLM generated data that may not be accurate, and it just con…

There must be signals in the data about generated garbage, otherwise humans wouldn't be able to tell. Something like PageRank would be a game changer and potentially solve this issue.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#29

Earlier quoted context omitted.

I can do it for $20B and a pony

Only because you're cheating by using the neurons of the pony.

Shh.

We're the only AI company that can offer HorseSense (TM)

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#30
post #15

LLMs are comprised of just three elements Data Compute Algorithms All three are just scratching the surface of what is possible. Data: What has been scraped off the internet is just Compute: While 3nm processes are approaching an atomic limit (0.21nm for Si), there is still room to explore more densely packed transistors or other materials like Gallium Nitride or optical computing. Not only that but there is a lot of…

For data though, as LLM's generate more output, over time wouldn't they be expected to mess themselves with their own generated data? Wouldn't that be the wall we'll hit? Think of how shitted up Google Search is with generated garbage, I'm imagining we're already in the 'golden age' where we were able to train on good datasets before it gets 'polluted' with LLM generated data that may not be accurate, and it just con…

Cleaning and preparing the dataset is a huge part of training. Like the OP mentioned, OpenAI likely have some high quality automation for doing this and that's what's given them a leg up above all other competitors. You can apply the same automation to clear out low quality AI content the same way you remove low quality human content. It's not about the source, just the quality matters.
Post reply on HN