Live data from Hacker News

No "Zero-Shot" Without Exponential Data

arxiv.org

61–70 of 123 posts

Re: No "Zero-Shot" Without Exponential Data

#61
post #21
post #18

My biggest worry is the idea of generating data as training data. We're obviously already unwittingly doing this, but once someone decides to augment low-volume segments of the dataset with generative input, we're going to start getting some really crappy feedback loops.

Do you ever have novel ideas while walking or in the shower? Well, you’re learning from data you generate.

https://arxiv.org/pdf/2305.17493

There's some cursory indication that in the long tail, training LLMs on LLM-generated data causes model collapse. Kind of like how if you photocopy a photocopy too many times the document becomes unreadable.

This isn't really surprising though. Neural networks at large are a form of lossy compression. You can't do lossy compression on artifacts recovered from lossy compression too many times. The losses stack.

Re: No "Zero-Shot" Without Exponential Data

#62
post #15

This feels like the worst possible outcome of the current AI hype. We've essentially been ripping off the entire internet and feeding it to the models already, spending many billions of dollars in the process. It's pretty much the largest possible dataset you can currently get, and due to the ever-increasing and now rapidly accelerated AI poisoning of the internet most likely the largest possible dataset which will e…

Quite crappy? It seems to me like the current SOTA is working plenty good enough for most use cases and like the highest impact (practical) way to improve right now is going to be advancements in domain knowledge acquisition and retention.

I sure haven't seen any of those "plenty good" results yet - current SOTA seems to be about as useful as semi-coherently gluing together random Google results. Good enough to perhaps replace a minimum-wage worker, not good enough to provide actual value when you care about the quality of the result.

This paper seems to suggest that significant advancement in domain knowledge acquisition and retention is exactly the problem, as you seem to need exponentially more data due to a lack of generalization. What's the point of a model which can perfectly quote Shakespeare if you're a programmer trying to refactor a proprietary codebase and it fails to make a link to whatever garbage it picked up from StackOverflow?

Re: No "Zero-Shot" Without Exponential Data

#63
post #49

Of course, it will require exponential data for zero shot. The keyword here is zero shot . If you think about it for a second, this applies to humans too. We also need exponential training data to do things without examples.

When we learn the grammar of our language, the teacher does not stand in front of the class and proceed to say a large corpus of examples of ungrammatical sentences, only the correct ones are in the training set.

When we learn to drive, we do not need to crash our car a thousand times in a row before we start to get it.

When we play a new board game for the first time, we can do it fairly competently (though not as good as experienced players) just by reading and understanding the rules.

Re: No "Zero-Shot" Without Exponential Data

#64

I've always been rather upset that it's fairly common to train on things like LAION or COCO and then "zero shot" test on ImageNet. Zero shot doesn't mean a held out set, it means disjoint classes. You can't train on all the animals in the zoo with sentences and then be surprised your model knows zebras. You need to train on horses and test on zebras.

Should we also try to teach children geometry and test them on calculus?

This illustrates two ways of teaching

I’ve experienced both, each at a different university

In one, professors would teach one thing then ask very different (and much harder) questions on tests

In the other, tests were more of a recap of the material up to that point

I definitely learned a lot more in the second case and was a lot more motivated. It also required more effort from the professors

The two methods also test different things. The recap one tests effort and dedication, if you do the work, you get the grade

The difficult tests measure either luck and/or creativity and problem solving under pressure. It’s not about doing the work, it’s about either being lucky or good at testing

Re: No "Zero-Shot" Without Exponential Data

#65
post #27
post #15

This feels like the worst possible outcome of the current AI hype. We've essentially been ripping off the entire internet and feeding it to the models already, spending many billions of dollars in the process. It's pretty much the largest possible dataset you can currently get, and due to the ever-increasing and now rapidly accelerated AI poisoning of the internet most likely the largest possible dataset which will e…

> not-entirely-useless but still quite crappy 15 months ago, general-purpose LLMs that have not been specifically trained on legal reasoning could score better than 90% of humans on the multistate bar exam, and these are humans who actually completed law school. General-purpose LLMs get similar results in medicine, and when the models are fine-tuned for medical diagnosis they're even better. And that was more than a…

I think this says more about the benchmark than the capabilities of the model. If it were the case that 90th percentile performance on the bar exam mean that a model was a 90th percentile lawyer, and we've had these models for 15 months (in fact longer), where are all the LLM lawyers? The lesson here is that a test designed for humans may not be equally representative of capabilities when given to an LLM.

Re: No "Zero-Shot" Without Exponential Data

#66
post #63
post #49

Of course, it will require exponential data for zero shot. The keyword here is zero shot . If you think about it for a second, this applies to humans too. We also need exponential training data to do things without examples.

When we learn the grammar of our language, the teacher does not stand in front of the class and proceed to say a large corpus of examples of ungrammatical sentences, only the correct ones are in the training set. When we learn to drive, we do not need to crash our car a thousand times in a row before we start to get it. When we play a new board game for the first time, we can do it fairly competently (though not as g…

please help yourself and do a quick Google search about "zero shot" and "few shot" learning.

Re: No "Zero-Shot" Without Exponential Data

#67

I've always been rather upset that it's fairly common to train on things like LAION or COCO and then "zero shot" test on ImageNet. Zero shot doesn't mean a held out set, it means disjoint classes. You can't train on all the animals in the zoo with sentences and then be surprised your model knows zebras. You need to train on horses and test on zebras.

Should we also try to teach children geometry and test them on calculus?

I think the authors are responding to a claim that AI is doing this: look, we taught them geometry, and now they know calculus! GP is saying that it's not a true zero-shot to have a separate test and training set, because the classes overlap. Similarly, the authors are saying "true zero-shot" is basically not happening, at least not nearly to the extent some are claiming.

So everyone here, including TFA, are all kinda doubting the same claim (our AI models can perform zero-shot generalizations) in different ways, I think?

Re: No "Zero-Shot" Without Exponential Data

#68
post #26

Earlier quoted context omitted.

I think that depends what your expectations are, and what you mean by another AI winter. We're just scratching the surface of what's possible with the current state of the art. Even if there are no major advances or breakthroughs in the near future, LLMs and associated technologies are already useful in many use cases. Or close enough to useful where engineering rather than science will be sufficient to overcome many…

"AI winter" was a political phenomenon that happened because the AI applications didn't hold water to the hype, and all of the gullible people that invested on them got burned out by the disparity. We are very clearly on that same path again. What leads to the conclusion that another winter is coming. But even the fact that people are talking about it is evidence it's not here yet, and as always with political phenom…

Except contemporary AI is useful on a day to day basis.

Re: No "Zero-Shot" Without Exponential Data

#69

Quite a few people saw this coming. It is still early to tell if we reached AI winter again or not, but at least we can see that news are slowing down.

Dunno exactly what you mean by ai winter, but the whole gen AI thing has kind of got all the attention and even if that route slows down there are a ton of other branches fruitful for development.

The issue, I think, is that if results don't follow expectations, given the very high costs, investors might get cold feet.

Sure we already have some interesting applications, but they are not exactly printing money.

Re: No "Zero-Shot" Without Exponential Data

#70
post #43

Earlier quoted context omitted.

Shouldn't there be enough known-good training content that we can use it to determine if a document is worth including in a training set?

There's not enough human workers to validate that at scale. If you want ML to do it... well that's a bit of a catch-22. How would an ML algorithm know if data is good enough to be trained on unless it has already been trained on that data?

You could use data where you know it was not AI generated, like the library of Congress catalog from prior to 2015. Or highly cited research papers, things of that nature.
Post reply on HN