Live data from Hacker News

No "Zero-Shot" Without Exponential Data

arxiv.org

31–40 of 123 posts

Re: No "Zero-Shot" Without Exponential Data

#31
post #21
post #18

My biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.

Do you ever have novel ideas while walking or in the shower? Well, you’re learning from data you generate.

I don't necessarily think it is the equivalent. Maybe it's more akin to a high-schooler reading a single book and then being asked to write several new books that hold equal weight in teaching the next generation of children.

These new text books could be great at simplifying the subject matter and making the material accessible or they may just never have fully understood the materials and are misleading.

Now imagine that over and over again, imo it's pretty likely to introduce inaccuracies if just taking a naiive approach.

Re: No "Zero-Shot" Without Exponential Data

#32
post #21
post #18

My biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.

Do you ever have novel ideas while walking or in the shower? Well, you’re learning from data you generate.

Well put.

Also, ever read your old journals? You are training on generated data.

Re: No "Zero-Shot" Without Exponential Data

#33
post #27
post #15

This feels like the worst possible outcome of the current AI hype. We've essentially been ripping off the entire internet and feeding it to the models already, spending many billions of dollars in the process. It's pretty much the largest possible dataset you can currently get, and due to the ever-increasing and now rapidly accelerated AI poisoning of the internet most likely the largest possible dataset which will e…

> not-entirely-useless but still quite crappy 15 months ago, general-purpose LLMs that have not been specifically trained on legal reasoning could score better than 90% of humans on the multistate bar exam, and these are humans who actually completed law school. General-purpose LLMs get similar results in medicine, and when the models are fine-tuned for medical diagnosis they're even better. And that was more than a…

It's not super surprising that LLMs perform well on standardized tests, given that they have a lot of standardized test related text in their training data. There are a lot of claims out there about the zero-shot ability of LLMs, and very little specific research to back it up. Until now that is.

Re: No "Zero-Shot" Without Exponential Data

#34
post #26

Quite a few people saw this coming. It is still early to tell if we reached AI winter again or not, but at least we can see that news are slowing down.

I think that depends what your expectations are, and what you mean by another AI winter. We're just scratching the surface of what's possible with the current state of the art. Even if there are no major advances or breakthroughs in the near future, LLMs and associated technologies are already useful in many use cases. Or close enough to useful where engineering rather than science will be sufficient to overcome many…

You're right, and I suspect the GP would agree with you about there being real engineering applications for LLMs, diffusion models, etc

But I think the term "AI Winter" usually refers to the underlying research programme and the economics around it. Soaking up many billions of dollars of industry money and public grants on the pitch that AGI might be just around the corner, and then being unable to deliver on that pitch can induce a hangover effect that makes it much harder to raise money for anything that smells even a little bit like the failed pitch. Investors and administrators feel burnt and turn very skeptical for a very long time.

Meanwhile, the actual productive applications which shook out of the initial boom just get renamed to something else so that they don't carry that smell.

We'll see how it goes here, but that's the familiar road and where the terminology comes from.

Re: No "Zero-Shot" Without Exponential Data

#35
post #18

My biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.

We are already doing this. In fact for the next generation of frontier models the primary electricity cost is running inference to generate training data for these models. As Llama 3 has shown us scale of training data is more important than size of model.

Re: No "Zero-Shot" Without Exponential Data

#36
post #27

Earlier quoted context omitted.

> not-entirely-useless but still quite crappy 15 months ago, general-purpose LLMs that have not been specifically trained on legal reasoning could score better than 90% of humans on the multistate bar exam, and these are humans who actually completed law school. General-purpose LLMs get similar results in medicine, and when the models are fine-tuned for medical diagnosis they're even better. And that was more than a…

It's not super surprising that LLMs perform well on standardized tests, given that they have a lot of standardized test related text in their training data. There are a lot of claims out there about the zero-shot ability of LLMs, and very little specific research to back it up. Until now that is.

This paper is about CLIP not LLMs and does not generalize to LLM architectures.

Re: No "Zero-Shot" Without Exponential Data

#37
post #28

Quite a few people saw this coming. It is still early to tell if we reached AI winter again or not, but at least we can see that news are slowing down.

> reached AI winter again Again?

It's a term for when AI hype dries up, has happened multiple times since the 60s[0].

[0]: https://en.wikipedia.org/wiki/AI_winter

Re: No "Zero-Shot" Without Exponential Data

#38
post #28

Quite a few people saw this coming. It is still early to tell if we reached AI winter again or not, but at least we can see that news are slowing down.

> reached AI winter again Again?

This was not the technology industry's first encounter with AI hype. The term was coined 40 years ago, and has been suggested as a description for almost a dozen periods in the field's history:

https://en.wikipedia.org/wiki/AI_winter

Re: No "Zero-Shot" Without Exponential Data

#39
post #18

My biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.

Shouldn't there be enough known-good training content that we can use it to determine if a document is worth including in a training set?

Re: No "Zero-Shot" Without Exponential Data

#40
post #21
post #18

My biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.

Do you ever have novel ideas while walking or in the shower? Well, you’re learning from data you generate.

[deleted]
Post reply on HN