Live data from Hacker News

LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

news.mit.edu

131–140 of 156 posts

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#131
post #51

This is amazing work, but to me it highlights some of the biggest problems in the current AI zeitgeist, we are not really trying to work on any neuron or ruleset that isnt much different from the perceptron thats just a sumnation function. Is it really that suprising that we just see this same structure repeated in the models. Just because feedforward topologies with single neuron steps are the easiest to train and r…

> Just because feedforward topologies with single neuron steps are the easiest to train and run on graphics cards does that really make them the actual best at accomplishing tasks? You are ignoring a mountain of papers trying all conceivable approaches to create models. It is evolution by selection, in the end transformers won.

> in the end transformers won

we're at the end?

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#132
post #125
post #107

Earlier quoted context omitted.

Evolution generally favors multiple winners in different roles over a single dominate strategy. People tend to favor single winners.

I both think this is a really astute and important observation and also think it's an observation that's more true locally than of people broadly. Modern neoliberal business culture generally and the consolidated current incarnation of the tech industry in particular have strong "tunnel vision" and belief in chasing optimality compared to many other cultures, both extant and past

In neoclassical economics, there are no local maxima, because it would make the math intractable and expose how much of a load of bullshit most of it is.

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#133
post #113

Earlier quoted context omitted.

> fairly solidly proven that LLMs aren't just lookup tables/stochastic parrots Well i'd strongly disagree. I see no evidence of this; I'm am quite well acquainted with the literature. All empirical statistical AI is just a means of approximating an empirical distribution. The problem with NLP is that there is no empirical function from text tokens to meanings; just as there is no function from sets of 2D images to a…

That suggests that no statistical method could ever recover hidden representations though. And that’s patently untrue. Taken to its greatest extreme you shouldn’t even be able to guess between two mixed distributions even when they have wildly non-overlapping ranges. Or put another way, all of statistical testing in science is flawed. I’m not saying you believe that, but I fail to see how that situation is structural…

Yes, I think most statistical testing in science is flawed.

But, to be clear, the reason it could ever work at all has nothing to do with the methods or the data itself, it has to do with the properties of the data generating process (ie., reality, ie., what's being measured).

You can never build representations from measurement data, this is called inductivism and it's pretty clearly false: no representation is obtained from just characterising measurement data. Theres no cases where I can think of that this would work -- temperature isnt patterns in thermometers; gravity isnt patterns in the positions of stars; and so on.

Rather you can decide between competing representations using stats in a few special cases. Stats never uncovers hidden representations, it can decide between different formal models which include such representations.

eg., if you characterise some system as having a power-law data generating process (eg., social network friendships), then you can measure some parameters of that process

or, eg., if you arrange all the data to already follow a law you know (eg., F=Gmm/r^2) then you can find G, 'statistically'.

This has caused a lot of confusion histroically: it seems G is 'induced over cases', but all the representaiton work has alerady been done. Stats/induction just plays the role of fine-tuning known representatios. it never builds any

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#134
post #53

Earlier quoted context omitted.

> have access to virtually the entire internet It isn't even close to 1% of the internet, much less virtually the entire internet. According to the latest dump, Common Crawl has 4.3B pages, but Google in 2016 estimated there are 130T pages. The difference between 130T and 4.3B is about 130T. Even if you narrow it down to Google's searchable text index it's "100's of billions of pages" and roughly 100P compared to Com…

130T unique pages? That seems highly unlikely as that averages to over 10000 pages for each human being alive. If gp merely wants texts of interest to self as opposed to an accurate snapshot it seems LLMs should be quite capable, one day.

Is it? Every user profile in every website is a page. Every single tweet is a page.

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#135

Earlier quoted context omitted.

While "local maximas" is wrong, I think "a local maxima" is a valid way to say "a member of the set of local maxima" regardless of the number of elements in the set. It could even be a singleton.

No, a member of the set of local maxima is a a local maximum, just like a member of the set of people is a person, because it is a definite singular. The plural is also used for indefinite number, so “the set of local maxima” remains correct even if the set has cardinality 1, but a member of the set has definite singular number irrespective of the cardinality of the set.

I've been convinced, thanks!

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#136

Llms seem like a good compression mechanism. It blows my mind that I can have a copy of llama locally on my PC and have access to virtually the entire internet

PAC learning is compression.

PAC learnable, Finite VC dimensionality, and the following form of compression are fully equivalent.

https://arxiv.org/abs/1610.03592

Basically each individual neuron/perceptron just splits a space into two subspaces.

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#138
post #42

Earlier quoted context omitted.

If you've read the article, the LLM hallucinations aren't due to the model not knowing the information but a function that choose to remember the wrong thing.

From the paper: > Finally, we use our dataset and LRE-estimating method to build a visualization tool we call an attribute lens. Instead of showing the next token distribution like Logit Lens (nostalgebraist, 2020) the attribute lens shows the object-token distribution at each layer for a given relation. This lets us visualize where and when the LM finishes retrieving knowledge about a specific relation, and can reve…

[deleted]

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#139
post #125

Earlier quoted context omitted.

I both think this is a really astute and important observation and also think it's an observation that's more true locally than of people broadly. Modern neoliberal business culture generally and the consolidated current incarnation of the tech industry in particular have strong "tunnel vision" and belief in chasing optimality compared to many other cultures, both extant and past

In neoclassical economics, there are no local maxima, because it would make the math intractable and expose how much of a load of bullshit most of it is.

Yep. This. It’s impressive how communication is instantaneous, unimpeded, complete and transparent in economics.

Those things aren’t even true in a 500 person company let alone an economy.

Re: LLMs use a surprisingly simple mechanism to retrieve some stored knowledge

#140
post #27

Earlier quoted context omitted.

I didn't imply that they know anything about where atoms are, I was just pointing out the sheer absurdity of that volume of data. I should make it clear that my comparison there is unfair and mostly just funny – you don't need to store every possible combination of 10 tokens, because most of them will be nonsense, so you wouldn't actually need that much storage. That being said, it's been fairly solidly proven that L…

> fairly solidly proven that LLMs aren't just lookup tables/stochastic parrots Well i'd strongly disagree. I see no evidence of this; I'm am quite well acquainted with the literature. All empirical statistical AI is just a means of approximating an empirical distribution. The problem with NLP is that there is no empirical function from text tokens to meanings; just as there is no function from sets of 2D images to a…

Too true. I often point out to others that a transformer like gpt-4 operates wholly on numbers- it knows nothing of meaning in the real world- nothing at all
Post reply on HN