Live data from Hacker News

Do Machine Learning Models Memorize or Generalize?

pair.withgoogle.com

71–80 of 217 posts

Re: Do Machine Learning Models Memorize or Generalize?

#71
post #46
post #32

Earlier quoted context omitted.

No doubt. But I have also seen what people thought were generalized models failing on outlier, but valid, data. Quite often. Put another way, it isn't just how simple this task seems to be in the number of terms that are important, but isn't it also a rather dense function? Probably better question to ask is how sensitive are models that are looking at less dense functions to this? (Or more dense.). I'm not trying to…

Maybe humans are also failing a lot in out of distribution settings. It might be inherent.

We have names for that. :D. Stereotypes being a large one. Racism being motivated interpretation on the same ideas. Right?

Re: Do Machine Learning Models Memorize or Generalize?

#72
post #18

Earlier quoted context omitted.

Generalize is seeing common principles, patterns, between disparate instances of a phenomena. It's a proper word for this.

That's a common mechanism to achieve generalization, but the term is a little more general (heh) than that. It specifically refers to correctly handling data that lives outside the distribution presented by the training data. It's a description of a behavior , not a mechanism. Which may or may not be appropriate depending on whether you are talking about *what* the model does or *how* it achieves it.

Kinda fuzzy what's "in the distribution", because it depends on how deeply the model interprets it. If it understands examples outside the distribution... that kinda puts them in the distribution.

General understanding makes the information in the distribution very wide. Shallow understanding makes it very narrow. Like say recognizing only specific combinations of pixels verbatim.

Re: Do Machine Learning Models Memorize or Generalize?

#73

Earlier quoted context omitted.

Whoever suggested 'eventual recovery from overfitting' is a kindred spirit. Why throw away the context and nuance? That decision only further leans into the 'AI is magic' attitude.

No, actually this is just how language evolves. I'm glad we have the word "car" instead of "carriage powered by internal combustion engine" even if it confused some people 100 years ago when the term became used exclusively to mean something a bit more specfic. Of course the jargon used in a specific sub-field evolves much more quickly than common usage because the intended audience of paper like this is expected to…

Language devolves just as it evolves. We (the grand we) regularly introduce ambiguity --words and meanings with no useful purpose, or that are worse than useless.

I'm not really weighing in on the appropriateness of the use "grok" in this case. It's just a pet peeve of mine that people bring out "language evolves" as an excuse for why any arbitrary change is natural and therefore acceptable and we should go with the flow. Some changes are strictly bad ones.

A go-to example is when "literally" no longer means "literally", but its opposite, or nothing at all. We don't have a replacement word, so now in some contexts people have to explain that they "literally mean literally".

Re: Do Machine Learning Models Memorize or Generalize?

#74

Sometimes I think the reason human memory in some sense is so amazing, is what we lack in storage capacity that machines have, we makeup for in our ability to create patterns that compress the amount of information stored dramatically, and then it is like we compress those patterns together with other patterns and are able to extract things from it. Like it is an incredibly lossy compression, but it gets the job done…

That’s not exactly true, there doesn’t seem to be an upper bound (that we can reach) on storage capacity in the brain [0]. Instead, the brain actually works to actively distill knowledge that doesn’t need to be memorized verbatim into its essential components in order to achieve exactly this “generalized intuition and understanding” to avoid overfitting. [0]: https://www.scientificamerican.com/article/new-estimate-bo…

Distilling knowledge is data compression.

Re: Do Machine Learning Models Memorize or Generalize?

#75

Sometimes I think the reason human memory in some sense is so amazing, is what we lack in storage capacity that machines have, we makeup for in our ability to create patterns that compress the amount of information stored dramatically, and then it is like we compress those patterns together with other patterns and are able to extract things from it. Like it is an incredibly lossy compression, but it gets the job done…

That’s not exactly true, there doesn’t seem to be an upper bound (that we can reach) on storage capacity in the brain [0]. Instead, the brain actually works to actively distill knowledge that doesn’t need to be memorized verbatim into its essential components in order to achieve exactly this “generalized intuition and understanding” to avoid overfitting. [0]: https://www.scientificamerican.com/article/new-estimate-bo…

I've thought about this a lot in the context of the desire people seem to have to try and achieve human immortality or at least indefinite lifespans. If SciAm is correct here and the upper bound is a quadrillion bytes, we may not be able to hit that given the bound on possible human experiences, but someone who lived long enough would eventually hit that. After a hundred million years or whatever the real number is of life, you'd either lose the ability to form new memories or you'd have to overwrite old ones to do so.

Aside from having to eventually experience the death of all stars and light and the decay of most of the universe's baryonic matter and then face an eternity of darkness with nothing to touch, it's yet another reason I don't think immortality (as opposed to just a very long lifespan) is actually desirable.

Re: Do Machine Learning Models Memorize or Generalize?

#76

Earlier quoted context omitted.

That is essentially what embeddings do

Maybe, except from my understanding an embedding vector tends to be much larger than the source token (due to the high dimensionality of the embedding space). So it's almost like a reverse compression in a way. That said I know vector DBs have much more efficient ways of storing those vector embedding.

Tokens are not 1:1 with vectors.

Re: Do Machine Learning Models Memorize or Generalize?

#77

Grr, the AI folks are ruining the term 'grok'. It means roughly 'to understand completely, fully'. To use the same term to describe generalization... just shows you didn't grok grokking.

I've always considered the important part of grokking something to be the intuitiveness of the understanding, rather than the completeness.

Re: Do Machine Learning Models Memorize or Generalize?

#78
post #18

Earlier quoted context omitted.

That's a common mechanism to achieve generalization, but the term is a little more general (heh) than that. It specifically refers to correctly handling data that lives outside the distribution presented by the training data. It's a description of a behavior , not a mechanism. Which may or may not be appropriate depending on whether you are talking about *what* the model does or *how* it achieves it.

Kinda fuzzy what's "in the distribution", because it depends on how deeply the model interprets it. If it understands examples outside the distribution... that kinda puts them in the distribution. General understanding makes the information in the distribution very wide. Shallow understanding makes it very narrow. Like say recognizing only specific combinations of pixels verbatim.

I think you are misinterpreting. The distribution present in the training set in isolation (the one I'm referring to, and is not fuzzy in the slightest) is not the same thing as the distribution understood by the trained model (the one you are referring to, and is definitely more conceptual and hard to characterize in non-trivial cases).

"Generalization" is simply the theoretical measure of how much the later extends beyond the former, regardless of how that's achieved.

Re: Do Machine Learning Models Memorize or Generalize?

#79

Sometimes I think the reason human memory in some sense is so amazing, is what we lack in storage capacity that machines have, we makeup for in our ability to create patterns that compress the amount of information stored dramatically, and then it is like we compress those patterns together with other patterns and are able to extract things from it. Like it is an incredibly lossy compression, but it gets the job done…

There are rare people who remember everything

https://youtu.be/hpTCZ-hO6iI

Re: Do Machine Learning Models Memorize or Generalize?

#80
post #73

Earlier quoted context omitted.

No, actually this is just how language evolves. I'm glad we have the word "car" instead of "carriage powered by internal combustion engine" even if it confused some people 100 years ago when the term became used exclusively to mean something a bit more specfic. Of course the jargon used in a specific sub-field evolves much more quickly than common usage because the intended audience of paper like this is expected to…

Language devolves just as it evolves. We (the grand we) regularly introduce ambiguity --words and meanings with no useful purpose, or that are worse than useless. I'm not really weighing in on the appropriateness of the use "grok" in this case. It's just a pet peeve of mine that people bring out "language evolves" as an excuse for why any arbitrary change is natural and therefore acceptable and we should go with the…

Language only evolves, "devolving" isn't a thing. All changes are arbitrary. Language is always messy, fluid and ambigious. You should go with the flow because being a prescriptivist about the way other people speak is obnoxious and pointless.

And "literally" has been used to mean "figuratively" for as long as the word has existed[0].

[0]https://blogs.illinois.edu/view/25/96439

Post reply on HN