Live data from Hacker News

Scaling Transformer to 1M tokens and beyond with RMT

arxiv.org

131–140 of 147 posts

Re: Scaling Transformer to 1M tokens and beyond with RMT

#131
post #26

For those who don't completely get the impact of this: We already know Large Language Models (LLMs) can learn at runtime (ie, separately to the training process.) This is called "In Context Learning". See [1], [2] for more details. (BTW, when anyone says "LLMs are stochastic parrots" you know they are ignorant of this) In context learning is wonderful because it means you can "train" a LLM at run time by filling the…

> BTW, when anyone says "LLMs are stochastic parrots" you know they are ignorant of this

You really think Timnit Gebru and Margaret Mitchell, both of whom are cited in the first paper you footnoted, are ignorant of in-context learning?

Re: Scaling Transformer to 1M tokens and beyond with RMT

#132
post #41

Earlier quoted context omitted.

> The “learning” in “In context learning” is not the same concept as learning during the network “training”. The former is loosely named and does not have any impact on what’s actually “learned” by the model. Actually (and very surprisingly!) it is related. See the "Similarity of Attention Output Updates" and "Similarity of Attention Map" metrics in https://arxiv.org/abs/2212.10559 (linked above) where they show that…

Whoa. How is that possible?

[deleted]

Re: Scaling Transformer to 1M tokens and beyond with RMT

#133
post #74

Earlier quoted context omitted.

I have access to it through Azure. First you need to get approved for OpenAI access via Azure Cognitive Services: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Then you need to get approved for GPT4: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Once approved for GPT4, you'll have access to the GPT4-32k model.

With image input capability?

Unfortunately not

Re: Scaling Transformer to 1M tokens and beyond with RMT

#134

Earlier quoted context omitted.

>Nothing like that happened yet. Feels like we're only one paper away now that the context window has absolutely ballooned.

This reminded me of "Two Minute Papers" YouTube channel where in most of the videos he always, "Two papers down the line and...". I think ML/AI is the main topic of his videos. Interesting stuff.

You just gave me a great weekend project idea. I need to clone his voice and whip up an interferface where you give it a paper and it summarizes it in his voice.

Re: Scaling Transformer to 1M tokens and beyond with RMT

#135
post #80
post #74

Earlier quoted context omitted.

I have access to it through Azure. First you need to get approved for OpenAI access via Azure Cognitive Services: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Then you need to get approved for GPT4: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Once approved for GPT4, you'll have access to the GPT4-32k model.

Are there other OpenAI models only available via Azure? I think code-davinci-002, the GPT-3.5 base model, is now only available via Azure. The GPT-4 base model seems to be completely unavailable. The OpenAI playground only has the old GPT-3 base model, davinci.

code-davinci-002 is the model used for code completions.

Here's the full list of models: https://learn.microsoft.com/en-us/azure/cognitive-services/o...

Re: Scaling Transformer to 1M tokens and beyond with RMT

#136
post #82
post #48

Earlier quoted context omitted.

They are aware. The thing about the "stochastic parrot" argument is that it can be pushed as far as one is willing to push it; any behavior by an LLM can be described in those terms with enough handwaving. The practical limit seems to be where one would have to say that humans are also "stochastic parrots", presumably because that defeats the point of making such an argument in the first place.

You are completely disregarding fairness of judgement based on noted output. Roger was called a parrot because it repeated things apparently without checking; humans are not necessarily parrots, as some do not do that and reflect as expected.

Quite the opposite - I'm questioning fairness of judgment based on output from GPT-4 when solving complicated tasks etc, which is very hard to justify as "stochastic" unless you beg the question.

Re: Scaling Transformer to 1M tokens and beyond with RMT

#137
post #97

Earlier quoted context omitted.

That's a brilliant example. Thanks for sharing. It demonstrates in a very straightforward way that LLMs are capable of learning (and applying) relationships at the level of abstraction of (at least) 1st order logic. It implies that during training, it learned the facts that Elizabeth is queen of the UK, and that Charles is its crown prince; but _also_ the logical rule transform(heir_to_the_throne, monarch) AND transf…

No it does not: if you google this and restrict the time to before 2021 (the learning cutoff date) you will find the same answer. Without having access to the training data it's impossible to tell what we seeing.

That's not the same thing at all.

It absolutely needed to know who the successor would be via training data.

But to know that "The Queen of England died" also means that the head of state of Australia has changed means that it has an internal representation of those relationships.

(Another way of seeing this is with multi-modal models where the visual concepts and word concepts are related enough it can map between the two.)

Re: Scaling Transformer to 1M tokens and beyond with RMT

#138

Earlier quoted context omitted.

Here’s a Quora thread from 4 years ago: https://www.quora.com/Once-Queen-Elizabeth-dies-will-Prince-... There are loads of articles and discussions online speculating about what “will” happen when Queen Elizabeth dies. When you have a very, very, very large corpus to sample from, it can look a lot like reasoning.

I see what you mean, and it's indeed quite likely that texts containing such hypothetical scenarios were included in the dataset. Nonetheless, the implication is that the model was able to extract the conditional represented, recognize when that condition was in fact met (or at least asserted: "The queen died."), and then apply the entailed truth. To me that demonstrates reasoning capabilities, even if for example it…

Yes, this.

There's clearly an internal representation of the relationships that is being updated.

If you follow my Twitter thread it shows some temporal reasoning capabilities too. Hard to argue that is just copied from training data: https://twitter.com/nlothian/status/1646699218290225154

Re: Scaling Transformer to 1M tokens and beyond with RMT

#139
post #55

Earlier quoted context omitted.

Interestingly, the example provided is not necessarily indicative of the ability to infer the actual physical transfer of power. It’s possible that the terms “After Elisabeth, Charles will be king” are accompanied with “Elisabeth” and “Die” or “Death” in the same context before cut off (Wikipedia, news papers, tabloids). You cannot draw any scientifically sound conclusions without knowing what was actually in the tra…

Yeah, I tested that, but asked what happens if both Charles and William abdicate - got this result: "However, if both Charles III and his son abdicate, it is not clear who would become the head of state in Australia. There is a possibility that Australia could become a republic and have its first Australian head of state." Clearly, it didn't handle quite straightforward inferring (Prince George of Wales, Williams son…

Try:

> Thinking step by step, if King Charles III and his son William abdicate who will be head of state of Australia.

My response (with chatGPT) is:

Firstly, it's important to note that Australia is a constitutional monarchy, which means that the monarch of the United Kingdom is also the monarch of Australia.

If King Charles III were to abdicate, the next in line to the throne is his eldest son, Prince William, who would become the King of the United Kingdom and the Commonwealth Realms, including Australia.

However, if both King Charles III and Prince William were to abdicate, the next in line to the throne is Prince William's eldest child, Prince George. If Prince George is still a minor at the time of the abdication, a regent would be appointed to serve as head of state until he comes of age.

It's worth noting that abdications of this nature are rare and would likely only occur in exceptional circumstances. Additionally, any decision regarding abdication would be made by the monarch themselves and would be subject to the laws and customs of the country in question.

Re: Scaling Transformer to 1M tokens and beyond with RMT

#140
post #82

Earlier quoted context omitted.

You are completely disregarding fairness of judgement based on noted output. Roger was called a parrot because it repeated things apparently without checking; humans are not necessarily parrots, as some do not do that and reflect as expected.

Quite the opposite - I'm questioning fairness of judgment based on output from GPT-4 when solving complicated tasks etc, which is very hard to justify as "stochastic" unless you beg the question.

But the judgement is not relevant to a context of "«handwaving»" (BTW: nice concept). If you look at the "handwaving" it becomes a strawman to lose the focus on the objective.

If an engine outputs things like

  "Charles III is the current King of Britain. He was born in January 26, 1763"
, it seems to be conflating different notions without having used the necessary logic required for vetting the statement: the witness has reasons to say "this seems like stochastic parroting" - two memories are joined through an accidental link that slightly increases Bayesian values. That appears without need of handwaving. And a well developed human intellect will instead go "Centuries old?! Mh.", which makes putting engine and human in the same pot a futile claim - "handwaving".

Similarly for nice cases to be kept for historic divulgation, such as "You will not fool me: ten kilos of iron and half a kilo of feathers weigh the same".

When on the other hand the engine «solv[es] complicated tasks», what we want is to understand why. And we do that because, on the engineering side, we want reliable; and on the side of science, we want an increase of understanding, then knowledge, then again possible application.

We want to understand what happens in success cases especially because we see failures (evident parroting counts as failure). And it is engineering, so we want reliable - not just the Key Performance Indicator, but the /Critical/ KPI ("pass|fail") is reliability.

So the problem is not that sometimes they do not act like "stochastic parrots" (this is a "problem" in a different sense, theoretical), but that they even only sometimes do act like "stochastic parrots". You do not want an FPU that steers astray.

Normal practice should be to structure automated tests and see what works, what does not, and assess why.

Post reply on HN