Live data from Hacker News

Language models are injective and hence invertible

arxiv.org

141–150 of 164 posts

Re: Language models are injective and hence invertible

#141
post #9

Earlier quoted context omitted.

Still, it is technically correct. The model produces a next-token likelihood distribution, then you apply a sampling strategy to produce a sequence of tokens.

Depends on your definition of the model. Most people would be pretty upset with the usual LLM providers if they drastically changed the sampling strategy for the worse and claimed to not have changed the model at all.

LLM providers are in the stone age with sampling today and it's on purpose because better sampling algorithms make the diversity of synthetic generated data too high, thus meaning your model is especially vulnerable to distillation attacks.

This is why we use top_p/top_k on the big 3 closed source models despite min_p and far better LLM sampling algorithms existing since 2023 (or in TFS case, since 2019)

Re: Language models are injective and hence invertible

#142
post #3

I don't like the title of this paper, since most people in this space probably think of language models not as producing a distribution (wrt which they are indeed invertible, which is what the paper claims) but as producing tokens (wrt which they are not invertible [0]). Also the author contribution statement made me laugh. [0] https://x.com/GladiaLab/status/1983812121713418606

But I bet you could reconstruct a plausible set of distributions by just rerunning the autoregression on a given text with the same model. You won't invert the exact prompt but it could give you a useful approximation.

Re: Language models are injective and hence invertible

#143

Author of related work here. This is very cool! I was hoping that they would try to invert layer by layer from the output to the input but it seems that they do a search process at the input layer instead. They rightly point out the residual connections make a layer by layer approach difficult. I may point out though that an rmsnorm layer should be invertible due to the epsilon term in the denominator which can be us…

What is meant by "residual connections" here?

Re: Language models are injective and hence invertible

#145
post #35
post #32

Earlier quoted context omitted.

Clarification [0] by the authors. In short: no, you can't. [0] https://x.com/GladiaLab/status/1983812121713418606

Thanks - seems like I'm not the only one who jumped to the wrong conclusion.

I also thought this when I read the abstract. input=prompt output=response does make more sense.

Re: Language models are injective and hence invertible

#146
post #37

>we confirm this result empirically through billions of collision tests on six state-of-the-art language models, and observe no collisions This sounds like a mistake. They used (among others) GPT2, which has pretty big space vectors. They also kind of arbitrarily define a collision threshold as an l2 distance smaller than 10^-6 for two vectors. Since the outputs are normalized, that corresponds to a ridiculously tiny…

As I read it, what they did there was a sanity-check by trusting the birthday paradox. Kind of: "If you get orthogonal vectors due to mere chance once, that's okay, but you try it billions of times and still get orthogonal vectors every time, mere chance seems a very unlikely explanation."

This has nothing to do with the birthday paradox. That paradox presumes a small countable state space (365) and a large enough # of observations.

In this case, it's a mathematical fact that 2 random vector in high dimensional space is very likely to be near orthogonal.

Re: Language models are injective and hence invertible

#147
post #24

I remember hearing an argument once that said LLMs must be capable of learning abstract ideas because the size of their weight model (typically GBs) is so much smaller than the size of their training data (typically TBs or PBs). So either the models are throwing away most of the training data, they are compressing the data beyond the known limits, or they are abstracting the data into more efficient forms. That's why…

> That's why an LLM (I tested this on Grok) can give you a summary of chapter 18 of Mary Shelley's Frankenstein, but cannot reproduce a paragraph from the same text verbatim. Unfortunately, the reality is more boring. https://www.litcharts.com/lit/frankenstein/chapter-18 https://www.cliffsnotes.com/literature/frankenstein/chapter-... https://www.sparknotes.com/lit/frankenstein/sparklets/ https://www.sparknotes.com/li…

I think you are probably right but it's hard to find an example of a piece of text that an LLM is willing to output verbatim (i.e. not subject to copyright guardrails) but also hasn't been widely studied and summarised by humans. Regardless, I think you could probably find many such examples especially if you had control of the LLM training process.

Re: Language models are injective and hence invertible

#148
post #117
post #24

I remember hearing an argument once that said LLMs must be capable of learning abstract ideas because the size of their weight model (typically GBs) is so much smaller than the size of their training data (typically TBs or PBs). So either the models are throwing away most of the training data, they are compressing the data beyond the known limits, or they are abstracting the data into more efficient forms. That's why…

There is a clarification tweet from the authors: - we cannot extract training data from the model using our method - LLMs are not injective w.r.t. the output text, that function is definitely non-injective and collisions occur all the time - for the same reasons, LLMs are not invertible from the output text https://x.com/GladiaLab/status/1983812121713418606

From the abstract:

> First, we prove mathematically that transformer language models mapping discrete input sequences to their corresponding sequence of continuous representations are injective

I think the "continuous representation" (perhaps the values of the weights during an inference pass through the network) is the part that implies they aren't talking about the output text, which by its nature is not a continuous representation.

They could have called out that they weren't referring to the output text in the abstract though.

Re: Language models are injective and hence invertible

#149
post #37

Earlier quoted context omitted.

As I read it, what they did there was a sanity-check by trusting the birthday paradox. Kind of: "If you get orthogonal vectors due to mere chance once, that's okay, but you try it billions of times and still get orthogonal vectors every time, mere chance seems a very unlikely explanation."

This has nothing to do with the birthday paradox. That paradox presumes a small countable state space (365) and a large enough # of observations. In this case, it's a mathematical fact that 2 random vector in high dimensional space is very likely to be near orthogonal.

A slightly stronger (and more relevant) statement is that the number of mutually nearly orthogonal vectors you can simultaneously pack into an N dimensional space is exponential in N. Here “mutually nearly orthogonal” can be formally defined as: choose some threshold epsilon>0 - the set S of unit vectors is nearly mutually orthogonal if the maximum of the pairwise dot products of between all members if S is less than epsilon. The statement of the exponential growth of the size of this set with N is (amazingly) independent of the value of epsilon (although the rate of growth does obviously depend on that value).

This is pretty unintuitive for us 3D beings.

Re: Language models are injective and hence invertible

#150
post #87

Earlier quoted context omitted.

I agree it is technically correct, but I still think it is the research paper equivalent of clickbait (and considering enough people misunderstood this for them to issue a semi-retraction that seems reasonable)

I disagree. Within the research community (which is the target of the paper), that title means something very precise and not at all clickbaity. It's irrelevant that the rest of the Internet has an inaccurate notion of "model" and other very specific terms.

In a field with as much public visibility as this one it is naive to only think of the academic target audience, especially when choosing a title like this. As a researcher you are responsible for communicating your findings both to other experts and to outsiders, and that includes choosing appropriate titles. (Though i think we fundamentally disagree about the role of researchers here) It's like writing a title that says "drinking only 200ml of water a day leads to weight loss" which is technically true, but misleading.
Post reply on HN