Live data from Hacker News

NanoGPT

github.com

51–60 of 334 posts

Re: NanoGPT

#51

Earlier quoted context omitted.

Yes, but to be fair, the system that does the training really sucks and doesn’t scale.

Neither does OpenAI. It costs so much and still delivers so little. A human can generate breakthroughs in science and tech that can be used to reduce carbon emissions. ChatGPT can do no such thing.

You can't know that. Currently, 8 billion humans generate a few scientific breakthroughs per year. You'd have to run several billion ChatGPTs for a year with zero breakthroughs to have any confidence in such a claim.

Re: NanoGPT

#52
post #33

Earlier quoted context omitted.

I can't find any source on the 1.5T params number. I'd love to read more if you have any links to share. Thanks

afaik, gpt-4 is mostly rumours so far, same thing for the 1.5T number. gpt-4 is suerly coming.

Maybe it will be called GPT-XP by then, with Microsoft owning half of it.

Re: NanoGPT

#53

Earlier quoted context omitted.

> Small players should focus on applications of this tech. That sounds a bit condescending. We are probably at a point from which the government should intervene and help establish level playing field. Otherwise we are going to see a deeper divide between multibillion businesses conquering multiple markets and sort of neofiefdom situation. This is not good.

It's not that condescending, that's todays reality. Should I feel entitled for $600k training time that may or may not work? Do you think the government is a good actor to judge if my qualifications are good enough to grant me resources worth a house? It's quite reasonable to make use of models already trained for small players.

> Do you think the government is a good actor to judge if my qualifications are good enough to grant me resources worth a house?

Governments already routinely do that for pharmaceutical research or for nuclear (fusion) research. In fact, almost all major impact research and development was funded by the government, mostly the military. Lasers, microwaves, silicon, interconnected computers - all funded by the US tax payer, back in the golden times when you'd get laughed out of the room if you dared think about "small government". And the sums involved were ridiculously larger than the worth of a house. We're talking of billions of dollars.

Nowadays, R&D funding is way WAY more complex. Some things like AI or mRNA vaccines are mostly funded by private venture capital, some are funded by large philanthropic donors (e.g. Gates Foundation), some by the inconceivably enormous university endowments, a lot by in-house researchers at large corporations, and a select few by government grants.

The result of that complexity:

- professors have to spend an absurd percentage of their time "chasing grants" (anecdata, up to 40% [1]) instead of actually doing research

- because grants are time-restricted, it's rare to have tenure track any more

- because of the time restriction and low grant amounts, it's very hard for the support staff as well. In Germany and Austria, for example, extremely low paid "chain contracts" are common - one contract after another, usually for a year, but sometimes as low as half a year. It's virtually impossible to have a social life if you have to up-root it for every contract because you have to take contracts wherever they are, and forget about starting a family because it's just so damn insecure. The only ones that can make it usually come from highly privileged environments: rich parents or, rarely, partners that can support you.

Everyone in academia outside of tenured professors struggles with surviving, and the system ruthlessly grinds people to their bones. It's a disgrace.

[1] https://www.johndcook.com/blog/2011/04/25/chasing-grants/

Re: NanoGPT

#54
post #41

Earlier quoted context omitted.

What does “355 years” mean in this context? I assume it’s not human years

Claimed here, so this is presumably the reference (355 GPU Years): https://lambdalabs.com/blog/demystifying-gpt-3 "We are waiting for OpenAI to reveal more details about the training infrastructure and model implementation. But to put things into perspective, GPT-3 175B model required 3.14E23 FLOPS of computing for training. Even at theoretical 28 TFLOPS for V100 and lowest 3 year reserved cloud pricing we could find…

That's still including margins of cloud vendors. OpenAI had Microsoft providing resources which could do that at much lower cost. It still won't be cheap but you'll be way below $5m if you buy hardware yourself, given that you're able to utilize it long enough. Especially if you set it up in a region with low electricity prices, latency doesn't matter anyway.

Re: NanoGPT

#55

Earlier quoted context omitted.

How long does it take to train a human? It's useless for two years then maybe it can tell you it needs to poop. The breakthrough will be developing this equivalent in an accessible manner and us taking care to train the thing for a couple of decades but then it becomes our friend.

Yes, but to be fair, the system that does the training really sucks and doesn’t scale.

I still think that this will be a major form of AI that is accessible to the public at large and it will enable productivity improvements at all levels.

I'm not joking, this is really something I think will/should happen.

Re: NanoGPT

#56

Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art

It should be no issue if it became massively parralelized a-la SETI. I wonder when Wikimedia or Apache foundation will jump into AI

Re: NanoGPT

#57

Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art

Alternatively, are there ways to train on consumer graphics cards, similar to SETI@Home or Folding@Home? I would personally be happy to donate gpu time, as I imagine many others would as well.

Re: NanoGPT

#58
post #56

Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art

It should be no issue if it became massively parralelized a-la SETI. I wonder when Wikimedia or Apache foundation will jump into AI

Wikimedia and other organizations that deal with moderation might want to keep this technology out of the hands of the general public for as long as possible.

Re: NanoGPT

#59

Are there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art

Could this be distributed? Put all those mining GPUs to work. A lot of people like participating in public projects like this. I would!

>> GPT-3 took 355 years to train

> Could this be distributed? Put all those mining GPUs to work.

Nope. It's a strictly O(n) process. If it weren't for the foresight of George Patrick Turnbull in 1668, we would not be anywhere close to these amazing results today.

Post reply on HN