Earlier quoted context omitted.
Yes, but to be fair, the system that does the training really sucks and doesn’t scale.
Neither does OpenAI. It costs so much and still delivers so little. A human can generate breakthroughs in science and tech that can be used to reduce carbon emissions. ChatGPT can do no such thing.
NanoGPT
51–60 of 334 posts
Re: NanoGPT
#52Earlier quoted context omitted.
I can't find any source on the 1.5T params number. I'd love to read more if you have any links to share. Thanks
afaik, gpt-4 is mostly rumours so far, same thing for the 1.5T number. gpt-4 is suerly coming.
Re: NanoGPT
#53Earlier quoted context omitted.
> Small players should focus on applications of this tech. That sounds a bit condescending. We are probably at a point from which the government should intervene and help establish level playing field. Otherwise we are going to see a deeper divide between multibillion businesses conquering multiple markets and sort of neofiefdom situation. This is not good.
It's not that condescending, that's todays reality. Should I feel entitled for $600k training time that may or may not work? Do you think the government is a good actor to judge if my qualifications are good enough to grant me resources worth a house? It's quite reasonable to make use of models already trained for small players.
Governments already routinely do that for pharmaceutical research or for nuclear (fusion) research. In fact, almost all major impact research and development was funded by the government, mostly the military. Lasers, microwaves, silicon, interconnected computers - all funded by the US tax payer, back in the golden times when you'd get laughed out of the room if you dared think about "small government". And the sums involved were ridiculously larger than the worth of a house. We're talking of billions of dollars.
Nowadays, R&D funding is way WAY more complex. Some things like AI or mRNA vaccines are mostly funded by private venture capital, some are funded by large philanthropic donors (e.g. Gates Foundation), some by the inconceivably enormous university endowments, a lot by in-house researchers at large corporations, and a select few by government grants.
The result of that complexity:
- professors have to spend an absurd percentage of their time "chasing grants" (anecdata, up to 40% [1]) instead of actually doing research
- because grants are time-restricted, it's rare to have tenure track any more
- because of the time restriction and low grant amounts, it's very hard for the support staff as well. In Germany and Austria, for example, extremely low paid "chain contracts" are common - one contract after another, usually for a year, but sometimes as low as half a year. It's virtually impossible to have a social life if you have to up-root it for every contract because you have to take contracts wherever they are, and forget about starting a family because it's just so damn insecure. The only ones that can make it usually come from highly privileged environments: rich parents or, rarely, partners that can support you.
Everyone in academia outside of tenured professors struggles with surviving, and the system ruthlessly grinds people to their bones. It's a disgrace.
[1] https://www.johndcook.com/blog/2011/04/25/chasing-grants/
Re: NanoGPT
#54Earlier quoted context omitted.
What does “355 years” mean in this context? I assume it’s not human years
Claimed here, so this is presumably the reference (355 GPU Years): https://lambdalabs.com/blog/demystifying-gpt-3 "We are waiting for OpenAI to reveal more details about the training infrastructure and model implementation. But to put things into perspective, GPT-3 175B model required 3.14E23 FLOPS of computing for training. Even at theoretical 28 TFLOPS for V100 and lowest 3 year reserved cloud pricing we could find…
Re: NanoGPT
#55Earlier quoted context omitted.
How long does it take to train a human? It's useless for two years then maybe it can tell you it needs to poop. The breakthrough will be developing this equivalent in an accessible manner and us taking care to train the thing for a couple of decades but then it becomes our friend.
Yes, but to be fair, the system that does the training really sucks and doesn’t scale.
I'm not joking, this is really something I think will/should happen.
Re: NanoGPT
#56Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art
Re: NanoGPT
#57Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art
Re: NanoGPT
#58Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art
It should be no issue if it became massively parralelized a-la SETI. I wonder when Wikimedia or Apache foundation will jump into AI
Re: NanoGPT
#59Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art
Could this be distributed? Put all those mining GPUs to work. A lot of people like participating in public projects like this. I would!
> Could this be distributed? Put all those mining GPUs to work.
Nope. It's a strictly O(n) process. If it weren't for the foresight of George Patrick Turnbull in 1668, we would not be anywhere close to these amazing results today.