Live data from Hacker News

GPT4 is 8 x 220B params = 1.7T params

twitter.com

101–110 of 215 posts

Re: GPT4 is 8 x 220B params = 1.7T params

#101

Earlier quoted context omitted.

How is it a debacle and how is it bad?

Because Twitter is trending towards bankruptcy. a) It is valued at about a 1/4 of what it was purchased at. b) Twitter Blue has generated an irrelevant amount of revenue and churn is increasing [1]. c) Roadmap looks poor. Video is a terrible direction where only Google, Amazon, TikTok etc have been able to make the numbers work and that's because it is subsidised through other revenue sources. Payments is DOA given T…

Twitter was trending towards bankruptcy anyway.

At best it has sped up towards that destination, at worst changes will eventually avoid it.

Re: GPT4 is 8 x 220B params = 1.7T params

#102
post #2

Is this still orders of magnitude smaller than a human brain? How many? Based on current human neurons/synapses knowledge?

We have no idea how to estimate the computational capacity of the brain at the moment. We can make silly estimates like saying that 1 human neuron is equivalent to something in an artificial network. But this is definitely wrong, biological neurons are far more complex than this. The big problem is that we don't understand the locus of computation in the brain. What is the thing performing the meaningful unit of comp…

This is an important point. On the one hand real neurons are a heck of a lot more complex than a single weight in a neural network, so exactly mimicking a human brain is still well outside our capabilities, even assuming we knew enough to build an accurate simulation of one. On the other hand, there's no intrinsic reason why you would need to in order to get similar capabilities on a lot of areas: especially when you consider that neurons have a very 'noisy' operation environment, it's very possible that there's a huge overhead to the work they do to make up for it.

Re: GPT4 is 8 x 220B params = 1.7T params

#103

Earlier quoted context omitted.

We really have no idea how to directly compare the two. Also, vast portions of the human brain are dedicated to the visual cortex, smelling, breathing, muscle control... things which have value to us but which don't contribute to knowledge work when evaluating how many parameters it would take to replace human knowledge work.

It's really interesting that human organism requires so much computational power to support live.

Not really, as we only use 10% of our brains. /s

Re: GPT4 is 8 x 220B params = 1.7T params

#104

« We can’t really make models bigger than 220B parameters » Can someone explains why?

Because of memory bandwidth. H100 has 3350gB/s of bandwidth, more gpus will give you more memory but not bandwidth. If you load 175b parameters in 8bit then you can get theoretically 3350/175=19 tokens/second. In MoE you need to process only one expert at a time so sparse 8x220b model would be only slightly slower than dense 220b model.

Okay, memory bandwidth certainly matters, but 19 tokens a second is not some fundamental lower limit on the speed of a language model and so this doesn't really explain why the limit would be 220b rather than say 440b or 800b?

Re: GPT4 is 8 x 220B params = 1.7T params

#105

Earlier quoted context omitted.

Because Twitter is trending towards bankruptcy. a) It is valued at about a 1/4 of what it was purchased at. b) Twitter Blue has generated an irrelevant amount of revenue and churn is increasing [1]. c) Roadmap looks poor. Video is a terrible direction where only Google, Amazon, TikTok etc have been able to make the numbers work and that's because it is subsidised through other revenue sources. Payments is DOA given T…

a) Meta had similar swings of over 3x between a 2021 high and a 2022 low. We can only guess what Twitter would be worth today if the stock was floated, but it isn’t. Substituting market valuations with our personal beliefs isn’t interesting. b) Every social media company has a well populated graveyard of failed experiments behind them. c) Opinion. d) Business as usual for every social media company since the dawn of…

a - f) Opinion.

Re: GPT4 is 8 x 220B params = 1.7T params

#106
post #80

I wouldn’t trust anything geohot says. He doesn’t have access to any inside information.

That's luckily not relevant here. Multiple sources have independently confirmed the same rumors. This means that there is a significant probability of it being true.

What exactly is "it"? What is being discussed here? The size of the model? Who cares? People will make the same API call to get the results. Everyone is acting like a glorified AI expert now smh

Re: GPT4 is 8 x 220B params = 1.7T params

#107
I often hear the idea of digital is faster then biology. This seems mostly derived from small math computations.

Yet it seems the current form of large language computations is much much slower then our biology. Making it even larger will be necessary to come closer to human levels but the speed?

If this is the path to GI, the computational levels need to be very High and very centralized.

Are there ways to improve this in its current implementation other the cache & more hardware?

Re: GPT4 is 8 x 220B params = 1.7T params

#108

I often hear the idea of digital is faster then biology. This seems mostly derived from small math computations. Yet it seems the current form of large language computations is much much slower then our biology. Making it even larger will be necessary to come closer to human levels but the speed? If this is the path to GI, the computational levels need to be very High and very centralized. Are there ways to improve t…

For all we know these models should be compressible to much smaller sizes.

Re: GPT4 is 8 x 220B params = 1.7T params

#109
post #74

Earlier quoted context omitted.

It's up to Geohot to prove his claim. Althoug he may have insider information, he isn't in a position known to have inside information. It might be more correct to phrase it that way but we all got the meaning.

That's not how it works out in practice. For example, journalists who have insider info from sources and leak it on the regular don't do anything to prove they have insider information, they build up reputation based on how many predictions they make. There's no point in outing your sources, that's how you lose them.

Until reputation is built you aren't trusted. It seems geohot's reputation still doesn't convince everyone.

Re: GPT4 is 8 x 220B params = 1.7T params

#110

Earlier quoted context omitted.

That's not how it works out in practice. For example, journalists who have insider info from sources and leak it on the regular don't do anything to prove they have insider information, they build up reputation based on how many predictions they make. There's no point in outing your sources, that's how you lose them.

Until reputation is built you aren't trusted. It seems geohot's reputation still doesn't convince everyone.

He has a reputation but it's for being untrustworthy.
Post reply on HN