Live data from Hacker News

OpenAI’s CEO says the age of giant AI models is already over

wired.com

11–20 of 525 posts

Re: OpenAI’s CEO says the age of giant AI models is already over

#11
What age? Like, 3 years?

On the other hand though, Chinchilla and multimodal approaches already showed how later AIs can be improved beyond throwing petabytes of data at them.

It is all about variety and quality from now on I think. You can teach a person all about the color zyra but without actually ever seeing it, they will never fully understand that color.

Re: OpenAI’s CEO says the age of giant AI models is already over

#13
post #12

I'm no expert but doesn't the architecture of minigpt4 that's on the front page right now give some indication of what the future might look like?

eh, I haven't personally found a usecase for LLMs yet given the fact that you can't trust the output and it needs to be verified by a human (which might as well be just as time consuming/expensive as actually doing the task yourself)

Re: OpenAI’s CEO says the age of giant AI models is already over

#16

What age? Like, 3 years? On the other hand though, Chinchilla and multimodal approaches already showed how later AIs can be improved beyond throwing petabytes of data at them. It is all about variety and quality from now on I think. You can teach a person all about the color zyra but without actually ever seeing it, they will never fully understand that color.

It does seem, though, that using chinchilla like techniques does not create a copy with the same quality as the original. It's pretty good for some definition of the phrase, but it isn't equivalent, it's a lossy technique.

Re: OpenAI’s CEO says the age of giant AI models is already over

#18
post #2

https://archive.is/s4V9e He did not say what kind of research strategies or techniques might take its place. In the paper describing GPT-4, OpenAI says its estimates suggest diminishing returns on scaling up model size. Altman said there are also physical limits to how many data centers the company can build and how quickly it can build them.

> In the paper describing GPT-4, OpenAI says its estimates suggest diminishing returns on scaling up model size.

I read the two papers (gpt 4 tech report, and sparks of agi) and in my opinion they don't support this conclusion. They don't even say how big GPT-4 is, because "Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar."

> Altman said there are also physical limits to how many data centers the company can build and how quickly it can build them.

OK so his argument is like "the giant robots won't be powerful, but we won't show how big our robots are, and besides, there are physical limits to how giant of a robot we can build and how quickly we can build it." I feel like this argument is sus.

Re: OpenAI’s CEO says the age of giant AI models is already over

#19

Once you've trained on the internet and most published books (and more...) what else is there to do? You can't scale up massively anymore.

They didn't train it on the entire internet tho, only a small amount (in comparison to entire internet). Still plenty they could do.

Re: OpenAI’s CEO says the age of giant AI models is already over

#20
Related reading: https://dynomight.net/scaling/

In short it seems like virtually all of the improvement in future AI models will come from better algorithms, with bigger and better data a distant second, and more parameters a distant third.

Of course, this claim is itself internally inconsistent in that it assumes that new algorithms won't alter the returns to scale from more data or parameters. Maybe a more precise set of claims would be (1) we're relatively close to the fundamental limits of transformers, i.e., we won't see another GPT-2-to-GPT-4-level jump with current algorithms; (2) almost all of the incremental improvements to transformers will require bigger or better-quality data (but won't necessarily require more parameters); and (3) all of this is specific to current models and goes out the window as soon as a non-transformer-based generative model approaches GPT-4 performance using a similar or lesser amount of compute.

Post reply on HN