Earlier quoted context omitted.
Because of memory bandwidth. H100 has 3350gB/s of bandwidth, more gpus will give you more memory but not bandwidth. If you load 175b parameters in 8bit then you can get theoretically 3350/175=19 tokens/second. In MoE you need to process only one expert at a time so sparse 8x220b model would be only slightly slower than dense 220b model.
Okay, memory bandwidth certainly matters, but 19 tokens a second is not some fundamental lower limit on the speed of a language model and so this doesn't really explain why the limit would be 220b rather than say 440b or 800b?
GPT4 is 8 x 220B params = 1.7T params
111–120 of 215 posts
Re: GPT4 is 8 x 220B params = 1.7T params
#112Earlier quoted context omitted.
He doesn't strike me as the type of person to lie (except when trolling). His reputation is solid enough that I'm sure he's had discussions with people in the space.
>his reputation is solid Eh, is it? Not sure if I consider him an authority on anything anymore. https://www.reddit.com/r/ProgrammerHumor/comments/z2y8i0/fro... >This is the interview. Build this feature. You don't get source access. Link the GitHub and license it MIT. is akin to "Build this for free, license it MIT so I can use it without any issues, and oh, btw, I dont have authority to hire you, teehee."
He was just clicking around compulsively, jumping between stuff randomly without even reading it, while not understanding what he was looking at. I have no idea how he's successful in tech, judging from what I saw there.
Re: GPT4 is 8 x 220B params = 1.7T params
#113I often hear the idea of digital is faster then biology. This seems mostly derived from small math computations. Yet it seems the current form of large language computations is much much slower then our biology. Making it even larger will be necessary to come closer to human levels but the speed? If this is the path to GI, the computational levels need to be very High and very centralized. Are there ways to improve t…
It might end up being the kind of thing where if you want to accurately model consciousness, you would need a computer the size of the universe, and it's gotta run for like 13.8 billion years or so, but that's just my own pure speculation - I don't think anybody even has a clue on where to start tackling this problem.
This is not to discourage progress in the field. I'm personally very curious to see where all this will lead. However, if I had to place a bet, it wouldn't be on GI coming any time soon.
Re: GPT4 is 8 x 220B params = 1.7T params
#114Earlier quoted context omitted.
I think frequency is not additive like the parameters. If we are looking for analogy in compute power, then FLOPs is a better analogy to parameters.
I'm fairly certain the point of the gp was that the number of parameters are also not additive.
just because the first selection layer is very thin doesn't mean that the network cannot be considered composable
Re: GPT4 is 8 x 220B params = 1.7T params
#115Earlier quoted context omitted.
To be fair, this was already a common whisper at the time, so it could be a Chinese whisper effect. Even I thought this on release day: https://news.ycombinator.com/item?id=35165874 What is weird is how competitive the open-source models are to the closed ones. For instance, PaLM-2 Bison is below multiple 12B models[0], despite it being plausibly much bigger[1]. The gap with GPT-4 is not that big; the best open-sourc…
From my perspective, there's a vast divide between open source models and GPT4 at present. The lmsys leaderboard rankings are derived from users independently engaging with the LLMs and opting for the answers they find most appealing. Consequently, the rankings are influenced not only by the type of questions users ask but also by their preference for succinctness in responses. When we venture into the realm of more…
Re: GPT4 is 8 x 220B params = 1.7T params
#116GPT-4 is 1.7T params in the same way that an AMD Ryzen 9 7950X is 72 GHz.
You might be surprised to learn that dell "hpc engineers" (or maybe HPE?) have attempted to sell me hardware using that exact logic to compare servers.
Given a highly-parallelizable numerics job, and all other things being equal, a 20 core 2.0Ghz machine will run the job in the same time as a 40 core 1.0Ghz machine.
...though usually we'd just give it in FLOPS - though FLOPS doesn't describe non-float operation performance.
Probably just a sales/marketing person who learned a new technical word and wants to sound knowledgable to the customer but it back-fires.
Re: GPT4 is 8 x 220B params = 1.7T params
#117Earlier quoted context omitted.
I'm fairly certain the point of the gp was that the number of parameters are also not additive.
if you have a basket with 4 apples and a basket with 3 pears, are you not having 7 fruits ? just because the first selection layer is very thin doesn't mean that the network cannot be considered composable
Re: GPT4 is 8 x 220B params = 1.7T params
#118Earlier quoted context omitted.
You might be surprised to learn that dell "hpc engineers" (or maybe HPE?) have attempted to sell me hardware using that exact logic to compare servers.
Well, they're not wrong... Given a highly-parallelizable numerics job, and all other things being equal, a 20 core 2.0Ghz machine will run the job in the same time as a 40 core 1.0Ghz machine. ...though usually we'd just give it in FLOPS - though FLOPS doesn't describe non-float operation performance. Probably just a sales/marketing person who learned a new technical word and wants to sound knowledgable to the custom…
Re: GPT4 is 8 x 220B params = 1.7T params
#119Earlier quoted context omitted.
We have no idea how to estimate the computational capacity of the brain at the moment. We can make silly estimates like saying that 1 human neuron is equivalent to something in an artificial network. But this is definitely wrong, biological neurons are far more complex than this. The big problem is that we don't understand the locus of computation in the brain. What is the thing performing the meaningful unit of comp…
This is an important point. On the one hand real neurons are a heck of a lot more complex than a single weight in a neural network, so exactly mimicking a human brain is still well outside our capabilities, even assuming we knew enough to build an accurate simulation of one. On the other hand, there's no intrinsic reason why you would need to in order to get similar capabilities on a lot of areas: especially when you…
Since it is already giving interesting results, let's say a "brain" is a connectome with its current information flow.
Comparing AI with a brain in terms of scale is somewhat hazardous but with what we know about real neurons and synapses, one brain is still several orders of magnitude above the current biggest AIs (not to mention, current AI is 2D and very local, as the brain is 3D and much less locality constrained). The "self-awareness" zone would need a at leas 1000x bigger, redondant, with a saveable flow of information, 3D with less locality, connectome that of the current biggest AI. Not to mention, realtime rich inputs/outputs, and years of training (like a baby human).
Ofc, this is beyond us, we have no idea of what's going on, and we probably won't. This is totally unpredictable, anybody saying otherwise is either trying to steal money for some BS AI research, or a the real genius.
Re: GPT4 is 8 x 220B params = 1.7T params
#120Earlier quoted context omitted.
a) Meta had similar swings of over 3x between a 2021 high and a 2022 low. We can only guess what Twitter would be worth today if the stock was floated, but it isn’t. Substituting market valuations with our personal beliefs isn’t interesting. b) Every social media company has a well populated graveyard of failed experiments behind them. c) Opinion. d) Business as usual for every social media company since the dawn of…
a - f) Opinion.