Live data from Hacker News

GPT4 is 8 x 220B params = 1.7T params

twitter.com

81–90 of 215 posts

Re: GPT4 is 8 x 220B params = 1.7T params

#81

I wouldn’t trust anything geohot says. He doesn’t have access to any inside information.

https://twitter.com/soumithchintala/status/16712671501017210... It looks like at least one other person has also heard the same information.

To be fair, this was already a common whisper at the time, so it could be a Chinese whisper effect. Even I thought this on release day: https://news.ycombinator.com/item?id=35165874

What is weird is how competitive the open-source models are to the closed ones. For instance, PaLM-2 Bison is below multiple 12B models[0], despite it being plausibly much bigger[1]. The gap with GPT-4 is not that big; the best open-source model on the leaderboard beats it 30% of the time.

[0]: https://chat.lmsys.org/?leaderboard

[1]: https://news.ycombinator.com/item?id=35893357

Re: GPT4 is 8 x 220B params = 1.7T params

#83
post #75
post #34

weird title, note that the tweet said "so yes, GPT4 is *technically* 10x the size of GPT3, and all the small circle big circle memes from January were actually... in the ballpark?" It's really 8 models that are 220B, which is not the same as one model that is 1.7T params. There have been 1T+ models via mixtures of experts for a while now. Note also the follow up tweet: "since MoE is So Hot Right Now, GLaM might be th…

(OP here) - yeah i know, but i also know how AI twitter works so I put both the headline and the caveats. i always hope to elevate the level of discourse by raising the relevant facts to those at my level/a little bit behind me in terms of understanding. think theres always a fine balance between getting deep/technical/precise and getting attention and you have to thread the needle in a way that feels authentic to yo…

Makes it worse, since you have this understanding and still went with this explanation that it's 1.7 trillion patients

Re: GPT4 is 8 x 220B params = 1.7T params

#85
post #71

GPT-4 is 1.7T params in the same way that an AMD Ryzen 9 7950X is 72 GHz.

I think frequency is not additive like the parameters. If we are looking for analogy in compute power, then FLOPs is a better analogy to parameters.

I'm fairly certain the point of the gp was that the number of parameters are also not additive.

Re: GPT4 is 8 x 220B params = 1.7T params

#86

Earlier quoted context omitted.

How is it a debacle and how is it bad?

Because Twitter is trending towards bankruptcy. a) It is valued at about a 1/4 of what it was purchased at. b) Twitter Blue has generated an irrelevant amount of revenue and churn is increasing [1]. c) Roadmap looks poor. Video is a terrible direction where only Google, Amazon, TikTok etc have been able to make the numbers work and that's because it is subsidised through other revenue sources. Payments is DOA given T…

a) Meta had similar swings of over 3x between a 2021 high and a 2022 low. We can only guess what Twitter would be worth today if the stock was floated, but it isn’t. Substituting market valuations with our personal beliefs isn’t interesting.

b) Every social media company has a well populated graveyard of failed experiments behind them.

c) Opinion.

d) Business as usual for every social media company since the dawn of time.

e) Opinion. And advertiser behaviour suggests otherwise.

f) History is littered with new entrants which fail to unseat the incumbent. It does happen, but it’s statistically rare.

Re: GPT4 is 8 x 220B params = 1.7T params

#87
post #74

Earlier quoted context omitted.

It’s fine to say you don’t believe him, it’s fine to say he’s wrong, but if you want to claim he doesn’t have inside information, I would expect a standard of evidence as good as that you would expect for Geohot’s claim. So tell me: how do you know the Geohot doesn’t have inside information?

It's up to Geohot to prove his claim. Althoug he may have insider information, he isn't in a position known to have inside information. It might be more correct to phrase it that way but we all got the meaning.

That's not how it works out in practice. For example, journalists who have insider info from sources and leak it on the regular don't do anything to prove they have insider information, they build up reputation based on how many predictions they make.

There's no point in outing your sources, that's how you lose them.

Re: GPT4 is 8 x 220B params = 1.7T params

#90

Earlier quoted context omitted.

https://twitter.com/soumithchintala/status/16712671501017210... It looks like at least one other person has also heard the same information.

To be fair, this was already a common whisper at the time, so it could be a Chinese whisper effect. Even I thought this on release day: https://news.ycombinator.com/item?id=35165874 What is weird is how competitive the open-source models are to the closed ones. For instance, PaLM-2 Bison is below multiple 12B models[0], despite it being plausibly much bigger[1]. The gap with GPT-4 is not that big; the best open-sourc…

From my perspective, there's a vast divide between open source models and GPT4 at present. The lmsys leaderboard rankings are derived from users independently engaging with the LLMs and opting for the answers they find most appealing. Consequently, the rankings are influenced not only by the type of questions users ask but also by their preference for succinctness in responses. When we venture into the realm of more complex tasks, such as code generation, GPT4 unequivocally outshines all open source models.

That said, the advancements in models like Orca and the "Textbooks are all you need" paper are noteworthy (https://arxiv.org/pdf/2306.02707.pdf, https://arxiv.org/abs/2306.11644). I'm optimistic about what future smaller models could achieve.

Post reply on HN