Live data from Hacker News

Neural Network Diffusion

arxiv.org

41–50 of 92 posts

Re: Neural Network Diffusion

#41

Earlier quoted context omitted.

> which is just a belief that magic is real Is there a law of thermodynamics which prevents AI from writing code which would train a better AI? Never learned that one in school. And FYI here's OpenAI plan to align superintelligence: "Our goal is to build a roughly human-level automated alignment researcher. We can then use vast amounts of compute to scale our efforts, and iteratively align superintelligence." I guess…

> Is there a law of thermodynamics which prevents AI from writing code which would train a better AI? You need to apply Wittgenstein here. This appears to be true because you haven't defined "better". If you define it, it'll become obvious that this is either false or true, but if it is true it'll be obvious in a way that doesn't make it sound interesting anymore. (For one thing our current "AI" don't come from "writ…

> they just come from training bigger models on the same data

Are you arguing that all AI models are using the same network structure?

This is only true in the most narrow sense, looking at models that are strictly improvements over previous generation models. It ignores the entire field of research that works by developing new models with new structures, or combining ideas from multiple previous works.

Re: Neural Network Diffusion

#42
post #27

Earlier quoted context omitted.

It's not magic, though. If AI can do work of a human, it can do work of a human. It's a trivial statement, and inability to see it is a hard cope. Are you gonna to take a bet "AI won't be able to do X in 10 years" for some X which people can learn to do now? If you're unwilling to bet then you believe that AI would plausibly be able to perform any human job, including job of AI researcher.

I don't claim it's impossible, just that there isn't a clear path from what exists now to that reality, and that the explanation presented by the above commenter (and I suppose OpenAI's website) does not clarify what they think the path is

What will AutoGPT look like if we have 100x more compute and another 10 years of research breakthroughs? It will be pretty damn good. If it can do the cognitive work of an AI researcher, well, there's your recursive self-improvement, at least on the research front (not so much on the hardware/energy front, physical constraints are trickier and will slow down progress in practice).

I don't know the exact path there, because if I did I'd publish and win the Turing Award. But it seems to be a plausible outcome in the medium-term future, at least if you go with Hinton's view that current methods are capable of understanding and reasoning, and not LeCun's view that it's all a dead end.

Re: Neural Network Diffusion

#43

Earlier quoted context omitted.

> Is there a law of thermodynamics which prevents AI from writing code which would train a better AI? You need to apply Wittgenstein here. This appears to be true because you haven't defined "better". If you define it, it'll become obvious that this is either false or true, but if it is true it'll be obvious in a way that doesn't make it sound interesting anymore. (For one thing our current "AI" don't come from "writ…

it is very clear to me that humans do in fact have a recursive self-improvement ability, and i'm confused why you think otherwise

A very small percentage maybe. I think I agree with the notion that most people bias toward thinking they are improving while actually self-sabotaging.

Re: Neural Network Diffusion

#44

Earlier quoted context omitted.

it is very clear to me that humans do in fact have a recursive self-improvement ability, and i'm confused why you think otherwise

I think people can read books (self improvement) and have children (recursive), but neither of those are both.

Why do you think that the human population is more intelligent, knowledgeable, and achieves greater technological feats as time goes on? It's because of recursive self-improvement, we are raised and educated into being better in a quite general sense, which includes being better at raising and educating; nearly every generation this cycle repeats and has for all of human history, at least since we acquired language. We also build machines that help us to make better machines, and then we use those better machines to make even better machines, another example of recursive self-improvement.

Re: Neural Network Diffusion

#45
post #31

Earlier quoted context omitted.

> which is just a belief that magic is real Is there a law of thermodynamics which prevents AI from writing code which would train a better AI? Never learned that one in school. And FYI here's OpenAI plan to align superintelligence: "Our goal is to build a roughly human-level automated alignment researcher. We can then use vast amounts of compute to scale our efforts, and iteratively align superintelligence." I guess…

Even if recursive self improvement does work out my hunch is that is going to be logarithmic instead of exponential mostly down to just availability of data. It might go beyond human intelligence but I don't think it will reach singularity

This is why the big bet for AI-assisted AI-development long term is synthetic data. A big part of the reason so much money and resources is going into synthetic data right now is not just out of economic necessity, but because there have been extremely encouraging results with synthetic data (e.g. 'Textbooks Are All You Need', AlphaZero).

Re: Neural Network Diffusion

#46

Earlier quoted context omitted.

it is very clear to me that humans do in fact have a recursive self-improvement ability, and i'm confused why you think otherwise

I think people can read books (self improvement) and have children (recursive), but neither of those are both.

New generations build onto the scientific knowledge of previous generations. It may not be fast but that sounds like recursive improvement to me. It seems reasonable for AI to accelerate this process.

Re: Neural Network Diffusion

#47

Earlier quoted context omitted.

> which is just a belief that magic is real Is there a law of thermodynamics which prevents AI from writing code which would train a better AI? Never learned that one in school. And FYI here's OpenAI plan to align superintelligence: "Our goal is to build a roughly human-level automated alignment researcher. We can then use vast amounts of compute to scale our efforts, and iteratively align superintelligence." I guess…

> Is there a law of thermodynamics which prevents AI from writing code which would train a better AI? You need to apply Wittgenstein here. This appears to be true because you haven't defined "better". If you define it, it'll become obvious that this is either false or true, but if it is true it'll be obvious in a way that doesn't make it sound interesting anymore. (For one thing our current "AI" don't come from "writ…

>> I guess people working there believe in magic.

>Yes, OpenAI was literally founded by a computer worshipping religious cult.

What cult is this?

Re: Neural Network Diffusion

#49
post #27

Earlier quoted context omitted.

I don't claim it's impossible, just that there isn't a clear path from what exists now to that reality, and that the explanation presented by the above commenter (and I suppose OpenAI's website) does not clarify what they think the path is

What will AutoGPT look like if we have 100x more compute and another 10 years of research breakthroughs? It will be pretty damn good. If it can do the cognitive work of an AI researcher, well, there's your recursive self-improvement, at least on the research front (not so much on the hardware/energy front, physical constraints are trickier and will slow down progress in practice). I don't know the exact path there, b…

I won't comment on whether I believe those researchers hold those views as you describe them, but as you describe them, I think both those descriptions of the state of AI research are untrue. The capabilities demonstrated by transformer models seem necessary but not sufficient to understand and reason, meaning that while they're not necessarily a "dead end", it is far from guaranteed that adding more compute will get them there

Of course if we allow for any arbitrary "research breakthrough" to happen then any outcome that's physically possible could happen, and I agree with you that superhuman artificial intelligence is possible. Nonetheless it remains unclear what research breakthroughs need to happen, how difficult they will be, and whether handing a company like OpenAI lots of money and chips will get that done, and it remains even more unclear whether that is a desirable outcome, given that the priorities of that company seem to shift considerably each time their budget is increased (As is the norm in this economic environment, to be clear, that is not a unique problem of OpenAI)

Obviously OpenAI has every reason to claim that it can do this and to claim that it will use the results in a way designed to benefit humanity as a whole. The people writing this promotional copy and the people working there may even believe both of these things. However, based on the information available, I don't think the first claim is credible. The second claim becomes less credible the more of the company's original mission gets jettisoned as its priorities align more to its benefactors, which we have seen happen rather rapidly

Re: Neural Network Diffusion

#50

I wasn't sure if this paper was parody on reading the abstract. It's not parody. Two things stand out to me: first is the idea of distilling these networks down into a smaller latent space, and then mucking around with that. That's interesting, and cross-sections a bunch of interesting topics like interpretability, compression, training, over- and under-.. The second is that they show the diffusion models don't just…

Seems like it could be useful for resizing the networks, no? Start with ChatGPT 4 then release an open version of it with much fewer parameters.

Or maybe some metaparameter that mucks with the sizes during training produces better results. Start large to get a baseline, then reduce size to increase coherence and learning speed, then scale up again once that is maxed out.

Post reply on HN