Live data from Hacker News

GPT4 is 8 x 220B params = 1.7T params

twitter.com

91–100 of 215 posts

Re: GPT4 is 8 x 220B params = 1.7T params

#94

Earlier quoted context omitted.

https://twitter.com/soumithchintala/status/16712671501017210... It looks like at least one other person has also heard the same information.

It’s funny that this post is trending on HN right next to the post about a paper showing how to build a model 1000x smaller than 1.7T that can code better than LLMs 10x larger.

I don't find it funny, I find it scary and mind-blowing: the impact of these headlines is additive - this one confirms the effectiveness of combining models, and the other one suggests you could cut the model size a couple orders of magnitude if you train on clean enough data. Together, this points at a way to achieve both GPT-4 that fits on your phone, and a much more powerful model that's not larger than GPT-4 is now.

Re: GPT4 is 8 x 220B params = 1.7T params

#95

I wouldn’t trust anything geohot says. He doesn’t have access to any inside information.

https://twitter.com/soumithchintala/status/16712671501017210... It looks like at least one other person has also heard the same information.

And from Bing - https://twitter.com/MParakhin/status/1670666605427298304

Re: GPT4 is 8 x 220B params = 1.7T params

#96

I wouldn’t trust anything geohot says. He doesn’t have access to any inside information.

It’s fine to say you don’t believe him, it’s fine to say he’s wrong, but if you want to claim he doesn’t have inside information, I would expect a standard of evidence as good as that you would expect for Geohot’s claim. So tell me: how do you know the Geohot doesn’t have inside information?

Prove a negative you mean?

Re: GPT4 is 8 x 220B params = 1.7T params

#97

Earlier quoted context omitted.

https://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh20... I think they _are_ going pretty extreme now.

Offtopic, but as a VR gamer that article just made me very sad. I was really hoping to see NVidia produce some decent cards in the near future, but looks like their main revenue is really going to be gargantuan number-crunchers. They'll likely only keep increasing the VRAM of gaming cards by arbitrarily-small numbers once every few years :-(

Gaming seems a lot less important than AI, in particular the graphical fidelity. Even games with crappy graphics can be fun. Crappy AI, not so much.

Re: GPT4 is 8 x 220B params = 1.7T params

#98

Earlier quoted context omitted.

It’s fine to say you don’t believe him, it’s fine to say he’s wrong, but if you want to claim he doesn’t have inside information, I would expect a standard of evidence as good as that you would expect for Geohot’s claim. So tell me: how do you know the Geohot doesn’t have inside information?

Prove a negative you mean?

It's up to anyone making a claim, if they want people who aren't yet convinced to become convinced, to provide some supporting evidence. This basic rule of discourse applies whether one makes a positive claim or a negative claim.

Bluedevilzn made a negative claim, so naturally I'm asking if they can prove their negative.

Re: GPT4 is 8 x 220B params = 1.7T params

#99
post #32

GPT-4 is 1.7T params in the same way that an AMD Ryzen 9 7950X is 72 GHz.

You might be surprised to learn that dell "hpc engineers" (or maybe HPE?) have attempted to sell me hardware using that exact logic to compare servers.

VMware still uses ghz for displaying load and compute capacity on your servers.

Re: GPT4 is 8 x 220B params = 1.7T params

#100

> GPT-4: 8 x 220B experts trained with different data/task distributions and 16-iter inference. There was a post on HackerNews the other day about a 13B open source model. Any 220B open source models? Why or why not? I wonder what the 8 categories were. I wonder what goes into identifying tokens and then trying to guess which category/model you should look up. What if tokens go between two models, how do the models r…

220B open source models wouldn't be as useful for most users.

You need two RTX 3090 24GB cards already to run inference with a 65B model that is 4bit quantized. Going beyond that (already expensive) hardware is out of reach for the average hobbyist developer.

Post reply on HN