Earlier quoted context omitted.
> GPT-4 recently extended that to 32,000 tokens Slightly unrelated, but has anyone outside of OpenAI got access to this model yet? While I received API access to GPT4 just a day or two after applying for it, I've yet to succeed to get access to the 32k version, nor has anyone I know, and I have not seen it being used by anyone in the wild either.
I have access to it through Azure. First you need to get approved for OpenAI access via Azure Cognitive Services: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Then you need to get approved for GPT4: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Once approved for GPT4, you'll have access to the GPT4-32k model.
Scaling Transformer to 1M tokens and beyond with RMT
81–90 of 147 posts
Re: Scaling Transformer to 1M tokens and beyond with RMT
#82Earlier quoted context omitted.
That isn't a fair summary of what I'm saying. People who say "LLMs are stochastic parrots" may well be aware of effects of conditioning the input. This in itself does not completely describe the capabilities of a LLM, since their ability to learn and use those new facts to override their previous "beliefs" is not what you'd expect from mere stochastic behavior. My point is that "stochastic parrot" believers are unawa…
They are aware. The thing about the "stochastic parrot" argument is that it can be pushed as far as one is willing to push it; any behavior by an LLM can be described in those terms with enough handwaving. The practical limit seems to be where one would have to say that humans are also "stochastic parrots", presumably because that defeats the point of making such an argument in the first place.
Re: Scaling Transformer to 1M tokens and beyond with RMT
#83Earlier quoted context omitted.
Or maybe it did. Who knows. If "both Charles III and his son abdicate" could well be considered indicative of some large upheaval or scandal, at which point it is entirely conceivable that the Australian electorate reaches a consensus on becoming a republic. The way that is phrased doesn't seem like a straightforward proposition to me at all.
I verified that it has all required facts (line of succession, current circumstances). I managed to get the right answer when got everything in context, but it failed again when all three abdicate (same context). Prince Harry was indicated once. I tested GPT a lot in other domains, what I found that as long the information explicitly exists (connection between facts) then the responses are fine. I assume that if GPT…
Feels like we're only one paper away now that the context window has absolutely ballooned.
Re: Scaling Transformer to 1M tokens and beyond with RMT
#84Earlier quoted context omitted.
I’m just saying, it was more a hopeful nod in the direction of wishful thinking. Not a joke, just a bit of having my head in the sand.
> I’m just saying, it was more a hopeful nod in the direction of wishful thinking. I still don’t understand why this is even wishful thinking, I don’t understand what the point of delaying progress is.
Re: Scaling Transformer to 1M tokens and beyond with RMT
#85For those who don't completely get the impact of this: We already know Large Language Models (LLMs) can learn at runtime (ie, separately to the training process.) This is called "In Context Learning". See [1], [2] for more details. (BTW, when anyone says "LLMs are stochastic parrots" you know they are ignorant of this) In context learning is wonderful because it means you can "train" a LLM at run time by filling the…
> "train" a LLM at run time by filling the context with examples.
Which is precisely what a stochastic parrot needs.
Re: Scaling Transformer to 1M tokens and beyond with RMT
#86Earlier quoted context omitted.
This is true too. Even chatGPT can learn some tool use in-context. See https://twitter.com/minosvasilias/status/1627076214639976449
GPT-4 writes some wicked complicated (but correct!) SQL if given a schema and a relevant task: https://gist.github.com/int19h/428cea1d87dfc389b99de1b79727f... What I found really amazing about this particular experiment is that the schema I gave it didn't contain any information that could be used to query for things like distances between places, and yet it came up with the idea of using settlements' culture as a pr…
Re: Scaling Transformer to 1M tokens and beyond with RMT
#87For those who don't completely get the impact of this: We already know Large Language Models (LLMs) can learn at runtime (ie, separately to the training process.) This is called "In Context Learning". See [1], [2] for more details. (BTW, when anyone says "LLMs are stochastic parrots" you know they are ignorant of this) In context learning is wonderful because it means you can "train" a LLM at run time by filling the…
> (BTW, when anyone says "LLMs are stochastic parrots" you know they are ignorant of this) Do you? Some of them -at least- understand that the output is conditional on the input.
I don't get what you are trying to say. That's like a property of any useful system.
Re: Scaling Transformer to 1M tokens and beyond with RMT
#88Earlier quoted context omitted.
> GPT-4 recently extended that to 32,000 tokens Slightly unrelated, but has anyone outside of OpenAI got access to this model yet? While I received API access to GPT4 just a day or two after applying for it, I've yet to succeed to get access to the 32k version, nor has anyone I know, and I have not seen it being used by anyone in the wild either.
I have access to it through Azure. First you need to get approved for OpenAI access via Azure Cognitive Services: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Then you need to get approved for GPT4: https://customervoice.microsoft.com/Pages/ResponsePage.aspx?... Once approved for GPT4, you'll have access to the GPT4-32k model.
Re: Scaling Transformer to 1M tokens and beyond with RMT
#89Earlier quoted context omitted.
I’m just saying, it was more a hopeful nod in the direction of wishful thinking. Not a joke, just a bit of having my head in the sand.
> I’m just saying, it was more a hopeful nod in the direction of wishful thinking. I still don’t understand why this is even wishful thinking, I don’t understand what the point of delaying progress is.
"Progress" doesn't mean "every massive change to the world and humanity is good." There are undeveloped technologies that we are currently not capable of being responsible with.
Google differential technological research.
Re: Scaling Transformer to 1M tokens and beyond with RMT
#90Earlier quoted context omitted.
That’s actually correct but an overfitted definition for learning. It holds certain hidden assumptions (i.e physical grounding) of the learner being human which makes it inapplicable to an LLM. As in a self driving car which passes a driving exam but fails to drive effectively freely in the city (it’s not an LLM but relevant in this context). You have to admit when you work with this tech that something fundamental i…
> That’s actually correct but an overfitted definition for learning. It holds certain hidden assumptions (i.e physical grounding) of the learner being human which makes it inapplicable to an LLM. Inapplicable why exactly? Because you say so? Logic isn't magic. Nor is learning. No (external) grounding is required either: iteratively eliminating inconsistent world models is all you need to converge toward a model of th…
Chill my friend, no need to get personal. We are talking about ideas. It’s OK to disagree. I am simply dismissing your initial claim. This usually happens when you present a scientific argument based on personal beliefs. If it’s not magic, then we should be able to doubt and examine it and it should eventually pass scientific muster.
> No grounding is required… It evidently managed to learn a pretty decent approximation.
Well, last time I used an LLM it suggested that I should lift the chair I am sitting in. I guess OpenAI has a lot of work to do. They have to eliminate this inconsistent world model for chairs, tables, floor, My dog, my cat and all the cats living on Mars…
edit: added a missing word.