Live data from Hacker News

GPT4 is 8 x 220B params = 1.7T params

twitter.com

11–20 of 215 posts

Re: GPT4 is 8 x 220B params = 1.7T params

#11

I wouldn’t trust anything geohot says. He doesn’t have access to any inside information.

He doesn't strike me as the type of person to lie (except when trolling). His reputation is solid enough that I'm sure he's had discussions with people in the space.

He's a crypto scammer. Look up cheap eth

Re: GPT4 is 8 x 220B params = 1.7T params

#12

Earlier quoted context omitted.

He doesn't strike me as the type of person to lie (except when trolling). His reputation is solid enough that I'm sure he's had discussions with people in the space.

He's a crypto scammer. Look up cheap eth

I mean that was marketed as a memecoin from the beginning. More so than doge even.

Re: GPT4 is 8 x 220B params = 1.7T params

#13

Earlier quoted context omitted.

It's really interesting that human organism requires so much computational power to support live.

I think it's even more interesting that the required amount of energy to do that high computational work isn't that high. Evolution has been working on it for a long time, and some things are really inefficient but overall it does an OK job at making squishy machines.

The human brain uses roughly 20 watts, which is really a remarkably low number.

https://psychology.stackexchange.com/a/12386

Re: GPT4 is 8 x 220B params = 1.7T params

#15
post #3
post #2

Is this still orders of magnitude smaller than a human brain? How many? Based on current human neurons/synapses knowledge?

2 orders of magnitude smaller, assuming 100T synaptic connections in the human brain.

It’s remarkable.

I’m curious how long until people are just using training brains in a jar to compute.

Re: GPT4 is 8 x 220B params = 1.7T params

#16

If it’s trivial then why does every other competitor suck at replicating it? Is it possible this is just a case of sour grapes that this intellectual is annoyed they’re not at the driving wheel of the coolest thing anymore?

I don't think he thinks its trivial - just that its not some revolutionary new architecture. It's an immense engineering effort rather than a fundamental breakthrough.

Re: GPT4 is 8 x 220B params = 1.7T params

#17
> GPT-4: 8 x 220B experts trained with different data/task distributions and 16-iter inference.

There was a post on HackerNews the other day about a 13B open source model.

Any 220B open source models? Why or why not?

I wonder what the 8 categories were. I wonder what goes into identifying tokens and then trying to guess which category/model you should look up. What if tokens go between two models, how do the models route between each other?

Re: GPT4 is 8 x 220B params = 1.7T params

#19
I really wonder if it is the case that the image processing is simply more tokens appended to the sequence. Would make the most sense from an architecture perspective, training must be a whole other ballgame of alchemy though

Re: GPT4 is 8 x 220B params = 1.7T params

#20
post #2

Is this still orders of magnitude smaller than a human brain? How many? Based on current human neurons/synapses knowledge?

We really have no idea how to directly compare the two. Also, vast portions of the human brain are dedicated to the visual cortex, smelling, breathing, muscle control... things which have value to us but which don't contribute to knowledge work when evaluating how many parameters it would take to replace human knowledge work.

While those portions of the brain aren't specific to learning intellectual or academic information, they might be crucial to making sense of data, help in testing what we learn, and help bridge countless gaps between model/simulation and reality (whatever that is). Hopefully that makes sense. Sort of like... Holistic learning.

I wonder if our brains and bodies are not all that separate, and the intangible features of that unity might be very difficult to quantify and replicate in silica.

Post reply on HN