> GPT-4: 8 x 220B experts trained with different data/task distributions and 16-iter inference. There was a post on HackerNews the other day about a 13B open source model. Any 220B open source models? Why or why not? I wonder what the 8 categories were. I wonder what goes into identifying tokens and then trying to guess which category/model you should look up. What if tokens go between two models, how do the models r…
GPT4 is 8 x 220B params = 1.7T params
21–30 of 215 posts
Re: GPT4 is 8 x 220B params = 1.7T params
#22Can someone explains why?
Re: GPT4 is 8 x 220B params = 1.7T params
#23I find it interesting that geohot says it is what you do “when you are out of ideas,” I can’t help but think that having multiple blended models is what makes GPT-4 seem like it has more “emergent” behavior than earlier models.
Re: GPT4 is 8 x 220B params = 1.7T params
#24I wouldn’t trust anything geohot says. He doesn’t have access to any inside information.
He doesn't strike me as the type of person to lie (except when trolling). His reputation is solid enough that I'm sure he's had discussions with people in the space.
Eh, is it? Not sure if I consider him an authority on anything anymore.
https://www.reddit.com/r/ProgrammerHumor/comments/z2y8i0/fro...
>This is the interview. Build this feature. You don't get source access. Link the GitHub and license it MIT.
is akin to "Build this for free, license it MIT so I can use it without any issues, and oh, btw, I dont have authority to hire you, teehee."
Re: GPT4 is 8 x 220B params = 1.7T params
#25« We can’t really make models bigger than 220B parameters » Can someone explains why?
Re: GPT4 is 8 x 220B params = 1.7T params
#26> GPT-4: 8 x 220B experts trained with different data/task distributions and 16-iter inference. There was a post on HackerNews the other day about a 13B open source model. Any 220B open source models? Why or why not? I wonder what the 8 categories were. I wonder what goes into identifying tokens and then trying to guess which category/model you should look up. What if tokens go between two models, how do the models r…
I think it’s just an ensemble of models, so you do some kind of pooling/majority vote on your output tokens
Re: GPT4 is 8 x 220B params = 1.7T params
#27Re: GPT4 is 8 x 220B params = 1.7T params
#28« We can’t really make models bigger than 220B parameters » Can someone explains why?
Re: GPT4 is 8 x 220B params = 1.7T params
#29Earlier quoted context omitted.
It doesn't fit in VRAM.
I’ve been a bit surprised that Nvidia hasn’t gone to extreme lengths to fit 1tb of memory on a card just for this reason.
I think they _are_ going pretty extreme now.
Re: GPT4 is 8 x 220B params = 1.7T params
#30Earlier quoted context omitted.
We really have no idea how to directly compare the two. Also, vast portions of the human brain are dedicated to the visual cortex, smelling, breathing, muscle control... things which have value to us but which don't contribute to knowledge work when evaluating how many parameters it would take to replace human knowledge work.
While those portions of the brain aren't specific to learning intellectual or academic information, they might be crucial to making sense of data, help in testing what we learn, and help bridge countless gaps between model/simulation and reality (whatever that is). Hopefully that makes sense. Sort of like... Holistic learning. I wonder if our brains and bodies are not all that separate, and the intangible features of…
"If our small minds, for some convenience, divide this glass of wine, this universe, into parts -- physics, biology, geology, astronomy, psychology, and so on -- remember that nature does not know it!" -Richard Feynmann