Live data from Hacker News

GPT4 is 8 x 220B params = 1.7T params

twitter.com

41–50 of 215 posts

Re: GPT4 is 8 x 220B params = 1.7T params

#41
post #19

I really wonder if it is the case that the image processing is simply more tokens appended to the sequence. Would make the most sense from an architecture perspective, training must be a whole other ballgame of alchemy though

Probably. Check the kosmos-1 paper from Microsoft that appeared a few days before GPT4 was released: https://arxiv.org/abs/2302.14045

Re: GPT4 is 8 x 220B params = 1.7T params

#43

> GPT-4: 8 x 220B experts trained with different data/task distributions and 16-iter inference. There was a post on HackerNews the other day about a 13B open source model. Any 220B open source models? Why or why not? I wonder what the 8 categories were. I wonder what goes into identifying tokens and then trying to guess which category/model you should look up. What if tokens go between two models, how do the models r…

same. i wish i had asked george instead of nodding along like an idiot. he probably wouldnt know but at least he’d speculate in interesting ways.

Re: GPT4 is 8 x 220B params = 1.7T params

#44

Earlier quoted context omitted.

He doesn't strike me as the type of person to lie (except when trolling). His reputation is solid enough that I'm sure he's had discussions with people in the space.

This wasn’t very impressive: https://www.pcmag.com/news/hacker-george-hotz-resigns-from-t... Spent two weeks trying to find someone to build a faceted search UI and then quit.

His biggest issue was trying to convince the remaining engineers and Elon that a refactor was more promising than continuing to build on the layers of spaghetti code that runs twitter.

I don't disagree with him, it's just clear that he didn't understand how dire the financial situation was and that even a progressive refactor starting with the most basic features would take considerable engineering hours and money.

Re: GPT4 is 8 x 220B params = 1.7T params

#45

Earlier quoted context omitted.

I’ve been a bit surprised that Nvidia hasn’t gone to extreme lengths to fit 1tb of memory on a card just for this reason.

https://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh20... I think they _are_ going pretty extreme now.

Offtopic, but as a VR gamer that article just made me very sad. I was really hoping to see NVidia produce some decent cards in the near future, but looks like their main revenue is really going to be gargantuan number-crunchers. They'll likely only keep increasing the VRAM of gaming cards by arbitrarily-small numbers once every few years :-(

Re: GPT4 is 8 x 220B params = 1.7T params

#46
post #3
post #2

Is this still orders of magnitude smaller than a human brain? How many? Based on current human neurons/synapses knowledge?

2 orders of magnitude smaller, assuming 100T synaptic connections in the human brain.

And how big is the memory and language part…?

Re: GPT4 is 8 x 220B params = 1.7T params

#47
post #30

Earlier quoted context omitted.

While those portions of the brain aren't specific to learning intellectual or academic information, they might be crucial to making sense of data, help in testing what we learn, and help bridge countless gaps between model/simulation and reality (whatever that is). Hopefully that makes sense. Sort of like... Holistic learning. I wonder if our brains and bodies are not all that separate, and the intangible features of…

We can say that such and such part of the brain is "for" this or that. Then it releases neurotransmitters or changes the level of hormones in your body which in turn have cascading effects, and at this point information theory would like to have a word. "If our small minds, for some convenience, divide this glass of wine, this universe, into parts -- physics, biology, geology, astronomy, psychology, and so on -- reme…

Reality only computes at the level of quarks - a less wrong post

Re: GPT4 is 8 x 220B params = 1.7T params

#48
post #39

Earlier quoted context omitted.

This wasn’t very impressive: https://www.pcmag.com/news/hacker-george-hotz-resigns-from-t... Spent two weeks trying to find someone to build a faceted search UI and then quit.

I think deciding to get away from the Musk/Twitter debacle as soon as you realize how bad it is isn't necessarily a bad thing..

How is it a debacle and how is it bad?

Re: GPT4 is 8 x 220B params = 1.7T params

#49

Are the models specifically trained to be experts in certain domains? Or the models are all trained on the same corpus, but just queried with different parameters? Is this functionally the same as beam search? Do they select the best output on a token-by-token basis, or do they let each model stream to completion and then pick the best final output?

Democracy of descendant models that have been trained separately by partitioning the identified clusters with strong capabilities from an ancestor model, so, in effect, they are modular, and can be learned to be combined competitively.

Re: GPT4 is 8 x 220B params = 1.7T params

#50

Earlier quoted context omitted.

https://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh20... I think they _are_ going pretty extreme now.

Offtopic, but as a VR gamer that article just made me very sad. I was really hoping to see NVidia produce some decent cards in the near future, but looks like their main revenue is really going to be gargantuan number-crunchers. They'll likely only keep increasing the VRAM of gaming cards by arbitrarily-small numbers once every few years :-(

[deleted]
Post reply on HN