I really wonder if it is the case that the image processing is simply more tokens appended to the sequence. Would make the most sense from an architecture perspective, training must be a whole other ballgame of alchemy though
GPT4 is 8 x 220B params = 1.7T params
41–50 of 215 posts
Re: GPT4 is 8 x 220B params = 1.7T params
#42Re: GPT4 is 8 x 220B params = 1.7T params
#43> GPT-4: 8 x 220B experts trained with different data/task distributions and 16-iter inference. There was a post on HackerNews the other day about a 13B open source model. Any 220B open source models? Why or why not? I wonder what the 8 categories were. I wonder what goes into identifying tokens and then trying to guess which category/model you should look up. What if tokens go between two models, how do the models r…
Re: GPT4 is 8 x 220B params = 1.7T params
#44Earlier quoted context omitted.
He doesn't strike me as the type of person to lie (except when trolling). His reputation is solid enough that I'm sure he's had discussions with people in the space.
This wasn’t very impressive: https://www.pcmag.com/news/hacker-george-hotz-resigns-from-t... Spent two weeks trying to find someone to build a faceted search UI and then quit.
I don't disagree with him, it's just clear that he didn't understand how dire the financial situation was and that even a progressive refactor starting with the most basic features would take considerable engineering hours and money.
Re: GPT4 is 8 x 220B params = 1.7T params
#45Earlier quoted context omitted.
I’ve been a bit surprised that Nvidia hasn’t gone to extreme lengths to fit 1tb of memory on a card just for this reason.
https://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh20... I think they _are_ going pretty extreme now.
Re: GPT4 is 8 x 220B params = 1.7T params
#46Re: GPT4 is 8 x 220B params = 1.7T params
#47Earlier quoted context omitted.
While those portions of the brain aren't specific to learning intellectual or academic information, they might be crucial to making sense of data, help in testing what we learn, and help bridge countless gaps between model/simulation and reality (whatever that is). Hopefully that makes sense. Sort of like... Holistic learning. I wonder if our brains and bodies are not all that separate, and the intangible features of…
We can say that such and such part of the brain is "for" this or that. Then it releases neurotransmitters or changes the level of hormones in your body which in turn have cascading effects, and at this point information theory would like to have a word. "If our small minds, for some convenience, divide this glass of wine, this universe, into parts -- physics, biology, geology, astronomy, psychology, and so on -- reme…
Re: GPT4 is 8 x 220B params = 1.7T params
#48Earlier quoted context omitted.
This wasn’t very impressive: https://www.pcmag.com/news/hacker-george-hotz-resigns-from-t... Spent two weeks trying to find someone to build a faceted search UI and then quit.
I think deciding to get away from the Musk/Twitter debacle as soon as you realize how bad it is isn't necessarily a bad thing..
Re: GPT4 is 8 x 220B params = 1.7T params
#49Are the models specifically trained to be experts in certain domains? Or the models are all trained on the same corpus, but just queried with different parameters? Is this functionally the same as beam search? Do they select the best output on a token-by-token basis, or do they let each model stream to completion and then pick the best final output?
Re: GPT4 is 8 x 220B params = 1.7T params
#50Earlier quoted context omitted.
https://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh20... I think they _are_ going pretty extreme now.
Offtopic, but as a VR gamer that article just made me very sad. I was really hoping to see NVidia produce some decent cards in the near future, but looks like their main revenue is really going to be gargantuan number-crunchers. They'll likely only keep increasing the VRAM of gaming cards by arbitrarily-small numbers once every few years :-(