Live data from Hacker News

Ask HN: Why is it taken for granted that LLM models will keep improving?

news.ycombinator.com

61–68 of 68 posts

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#61
post #35

Your skepticism is, I think, very well founded -- especially with such unclear definitions of "improvement." I think I have a corollary type idea: Why are LLM's not perhaps like "Linux," something than never really needs to be REWRITTEN from scratch, merely added to or improved on? In other words, isn't it fair to think that LoRA's are the really important thing to pay attention to? (And perhaps, like Google Fuschia…

When in comes to training LLMs, the definition of “improvement” is incredibly clear, as one must literally code a loss function that the model then minimizes.

It gets murkier trying to map that actual capabilities, but so far, lower loss has led to much stronger capabilities.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#62
post #59
post #25

Earlier quoted context omitted.

I think $80B is very cheap because the outcome is going to be a technological utopia. Honestly, I think my price is a bargain deal for the inhabitants of Earth.

> I think $80B is very cheap because the outcome is going to be a technological utopia. I wish I shared your optimism. However I've seen no evidence that society is prepared to deal with a large swath of jobs being obsoleted by AI. I have no doubt that the "haves" will call it a technological utopia, but I strongly suspect the "have nots" will be larger than ever.

In my architecture everyone is treated equally like an idiot so there are no haves and have not because the machine god treats all people the same and provides for all their needs within the "panoptic computronium cathedral"™.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#63
post #6

Because a lot of smart people are spending a lot of time, money, and effort on this. It's as simple as that. We could go into all sorts of details, like how increase in GPU capabilities will improve training capabilities, both in size and speed, or how GPU(/TPU) capabilities will improve, or how better techniques will make training on the same data set result in better models, or where other improvements will make be…

I mean a lot of highly payed/intelligent people worked in crypto/fusion/quantum computing, still all of those topics are evolving rather gradually.

Crypto is actually picking up after FTX and Binance, though as it's a social problem of adoption rather than a technical issue, I'm not sure it's comparable.

The problem with fusion and quantum computing is that advances are being made, but because those advances aren't consumer facing, you don't see them. Eg December 2022, they managed to get more energy out of a fusion experiment than they put in. That's huge! I'm not going to see an effect on my power bill for another couple decades, if ever, but it's real actual solid progress. For quantum computing, they're moving past the singular q-bit tech demonstrations level and moving into actual practical applications like making chips that can talk to each other **. Again, doesn't remotely affect me or my laptop today, but we've moved past the 1998 Stanford/IBM 2 q-bit computer.

Meanwhile, I can adopt a new model getting dropped with an afternoon of work, and see the results in milliseconds, in the case of StableDiffusion-turbo.

* https://www.technologyreview.com/2023/11/16/1083491/whats-co... ** https://www.technologyreview.com/2023/01/06/1066317/whats-ne...

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#64
post #58

Earlier quoted context omitted.

We need models that need less language data to train. Babies learn to talk on way less data than the entire internet. We need something closer to human experience. Kids have a feel for what is bullshit before they have consumed the entire internet :-). I think feeding the internet into a LLM will be seen as the mainframe days of AI.

> Babies learn to talk on way less data than the entire internet. Is this actually true? My gut check says yes, but I'm also unaware of any meaningful way to actually quantify the volume of sensor data processed by a baby (or anyone else for that matter), and it wouldn't shock me to discover if we could we'd find it to be a huge volume.

Babies in ancient societies certainly had less exposure to written language, much lower vocabulary, less exposure to music, etc.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#65
post #58

Earlier quoted context omitted.

We need models that need less language data to train. Babies learn to talk on way less data than the entire internet. We need something closer to human experience. Kids have a feel for what is bullshit before they have consumed the entire internet :-). I think feeding the internet into a LLM will be seen as the mainframe days of AI.

> Babies learn to talk on way less data than the entire internet. Is this actually true? My gut check says yes, but I'm also unaware of any meaningful way to actually quantify the volume of sensor data processed by a baby (or anyone else for that matter), and it wouldn't shock me to discover if we could we'd find it to be a huge volume.

Ah yes. I should be more precise. Less data that is textual. Of course other data sources are plentiful. Including internal and external sensory.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#66
post #58

Earlier quoted context omitted.

> Babies learn to talk on way less data than the entire internet. Is this actually true? My gut check says yes, but I'm also unaware of any meaningful way to actually quantify the volume of sensor data processed by a baby (or anyone else for that matter), and it wouldn't shock me to discover if we could we'd find it to be a huge volume.

Babies in ancient societies certainly had less exposure to written language, much lower vocabulary, less exposure to music, etc.

Sure the breadth is (maybe) smaller, but the question is volume. Babies get years of people talking around them, as well as data from their own muscles and vocalizations fed back to them. Is the volume they have consumed to the point the begin talking actually less than the volume consumed by an LLM?

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#67
post #66

Earlier quoted context omitted.

Babies in ancient societies certainly had less exposure to written language, much lower vocabulary, less exposure to music, etc.

Sure the breadth is (maybe) smaller, but the question is volume. Babies get years of people talking around them, as well as data from their own muscles and vocalizations fed back to them. Is the volume they have consumed to the point the begin talking actually less than the volume consumed by an LLM?

If you’re taking about babies in ancient societies (which I am), the answer is absolutely yes. They were exposed to much less language, and much less sound, than we are.

Re: Ask HN: Why is it taken for granted that LLM models will keep improving?

#68
post #66

Earlier quoted context omitted.

Sure the breadth is (maybe) smaller, but the question is volume. Babies get years of people talking around them, as well as data from their own muscles and vocalizations fed back to them. Is the volume they have consumed to the point the begin talking actually less than the volume consumed by an LLM?

If you’re taking about babies in ancient societies (which I am), the answer is absolutely yes. They were exposed to much less language, and much less sound, than we are.

Really? How much less? I'm far from convinced that if you sum up the sheer volume of noises heard, as well the other neurological inputs that goes into learning to speak (ex proprioception) you'd come out with a lesser number than what LLMs are trained on, but I'm open to any real data on this.
Post reply on HN