GPT-4 Architecture
twitter.com
GPT-4 Architecture
1–10 of 15 posts
Re: GPT-4 Architecture
#2Re: GPT-4 Architecture
#3Re: GPT-4 Architecture
#4A long intro with no real content, just tech bro "here's the thing" stuff trying to bait you into subscribing. Doesn't actually explain the architecture if you don't subscribe. Don't bother reading.
Re: GPT-4 Architecture
#5A long intro with no real content, just tech bro "here's the thing" stuff trying to bait you into subscribing. Doesn't actually explain the architecture if you don't subscribe. Don't bother reading.
The info is on the twitter link.
"GPT-4's details are leaked.
It is over.
Everything is here:"
Re: GPT-4 Architecture
#6The reason companies/researchers haven't generally touched MoE for LLMs despite how good it sounds on paper is because they've typically sucked and underperformed their dense counterparts.
assuming this is all true, Did Open ai do anything differently here or is it just scale ?
I know this very recent paper shows MoE benefit far more from Instruct tuning - https://arxiv.org/abs/2305.14705
FLAN-MOE-32B comfortably surpasses FLAN-PALM-62B with a third of the compute. It goes from 25.5% to 65.4% on MMLU.
In comparison, 55.1 to 59.6% for Flan-Palm 62b. That just kind of shows the underperformance you expect from sparse models.
But from Open ai's technical report, it doesn't seem like they needed that.
The Vision component seems to be just scale. Well all of it seems to be just scale. Seems like there's plenty scale left too as far as performance gains go.
Re: GPT-4 Architecture
#7Earlier quoted context omitted.
The info is on the twitter link.
I don't have a twitter account, I can only see the top level tweet that says "GPT-4's details are leaked. It is over. Everything is here:"
Re: GPT-4 Architecture
#8Re: GPT-4 Architecture
#9Re: GPT-4 Architecture
#10Earlier quoted context omitted.
Thanks!
not found - is there an archived version I can take a look at?