Live data from Hacker News

DreamFusion: Text-to-3D using 2D Diffusion

dreamfusion3d.github.io

191–200 of 208 posts

Re: DreamFusion: Text-to-3D using 2D Diffusion

#191

So does this mean I can use DreamBooth to create plausible NERFs of myself in any scenario? The future is looking weird.

Nah. This is made by Google, so it'll never become useful. You'll have to wait a few months for someone else to replicate it.

Someone did implement DreamBooth for stable diffusion so I was imagining, like you say, this (implemented by someone else) in a couple months with DreamBooth + stable diffusion

Re: DreamFusion: Text-to-3D using 2D Diffusion

#192
post #63

Earlier quoted context omitted.

Time and time again these ML techniques are proving to be wildly modular and pluggable. Maybe sooner or later someone will build a framework for end to end text-to-effective-ML-architecture that will just plug different things together and optimize them.

I think this is what huggingface (github for machine learning) is trying with diffusers lib: https://huggingface.co/docs/diffusers/index They have others as well.

Cool stuff. But who is working on the text-to-ML-architecture thing?

Re: DreamFusion: Text-to-3D using 2D Diffusion

#193
post #63

Earlier quoted context omitted.

Time and time again these ML techniques are proving to be wildly modular and pluggable. Maybe sooner or later someone will build a framework for end to end text-to-effective-ML-architecture that will just plug different things together and optimize them.

I think this is what huggingface (github for machine learning) is trying with diffusers lib: https://huggingface.co/docs/diffusers/index They have others as well.

Fascinating stuff! But who is working on the text-to-ML-architecture thing?

Re: DreamFusion: Text-to-3D using 2D Diffusion

#194
post #57

Earlier quoted context omitted.

> This seems like basically plugging a couple of techniques together that already existed [...] In his Lex Fridman interview, John Carmack makes similar assertions about this prospect for AGI: That it will likely be the clever combination of existing primitives (plus maybe a couple novel new ones) that make the first AGI feasible in just a couple thousand lines of code.

Billions of creatures with stronger neutral networks, more parameters, better input have lived on earth for millions of years, but only now something like humans showed up. I fully expect AI to do everything animals can do pretty soon, but since whatever it is that differentiates humans didn't happen for million of years, there's good chance AGI research will get stuck at a similar point.

Change arrives gradually, and then suddenly.

It takes nature thousands of years to create a rock that looks like a face, just by using geology. A human can do that in a couple hours. And then this AI can generate 50 3d human faces per second (assuming enough CPU).

It could be that an AGI is around the corner, as they say. We might not be machines, but are way faster than nature at reaching places. We don't have the option of waiting for thousands of years.

Re: DreamFusion: Text-to-3D using 2D Diffusion

#195
post #51

The thing that frightens me is that we are rapidly reaching broad humanity disrupting ML technologies without any of the social or societal frameworks to cope with it.

I'm usually not a fan of this general hand wringing / fear mongering around ML that a lot of people with too much time and not enough STEM background constantly bring up. Stable diffusion has been made available to the public for quite a while now and if anything has disproved a lot of the ungrounded nonsense that made companies like OpenAI censor their generative models.

I don't think you understand.

I watched this humor group when I was little (think 30 years ago). Usually they are inoffensive, but they have this particular sketch were a (man crossdressed as a) woman complains about "his husband hitting her" ... and you can see that one of her eyes is black. With canned laughs during the whole thing.

This used to make people laugh. 30 years ago. Then we had social changes and gradually our perspective as a society shifted, and this humor piece looks ... grotesque.

That's how I interpret the "societal changes" the OP is referring to. We have shiny new things every week and there's no time to adapt to them.

Re: DreamFusion: Text-to-3D using 2D Diffusion

#197

Earlier quoted context omitted.

You just picked an arbitrary gap though. And it seems like a small gap that's crossable. For example just 2 weeks ago there were two other gaps that you could've used in your example. You could've said 3D interpretation of images wasn't possible and the creation of animated movies wasn't possible and you could've said these few things suggest that there's a gap in what the human brain is doing and what the AI is doin…

I'm saying something different, that the most impressive examples of AI breakthroughs are doing things that humans find hard / are bad at. Meanwhile there are many things that people find easy / do without thinking / can be done by dogs or very young children that AI struggles with. It suggests to me that what most current approaches are doing is something fairly different from what human / animal intelligence is doi…

you will find that GPT-3 writing is consistently on all fronts better than what a 4 year old can write. Even the strange inconsistencies and lack of awareness in the writing is better than what a 4 year old can do.

Re: DreamFusion: Text-to-3D using 2D Diffusion

#198
post #40

Earlier quoted context omitted.

Why the downvote? I wasn't being sarcastic, it was a honest question, I'm really impressed how far this technology has come since GPT-3 2 years ago to DALl-E and Stable Diffusion ro Meta's text to video to this...

Maybe because you said "Metaverse" (and to some extent "VR") making it sound like sci-fi nonsense. You could have just said: How long then until we get photorealistic AI generated 3D games and experiences?

3D game and VR game (i.e. stereoscopic 3D game) is not the same experience.

But I agree that particular company is creating a bad marketing around the term Metaverse. (But pretty decent HW.) NVIDIA has much better footing with their Omniverse. But in the end, we all know the thing will be build as web browsers are today. USD + JS + WebRTC + WebXR can go a long way.

Re: DreamFusion: Text-to-3D using 2D Diffusion

#199

Earlier quoted context omitted.

>cat wearing sunglasses, back view") Bad prompt, missing implied antecedent/ambiguous subject... You may want: Back view of a cat which is wearing sunglasses, back view of a cat, but the view is wearing sunglasses, etc... I actually tried using projective terms from drafting books, and didn't get great results. Nor anatomicals either.

>Back view of a cat which is wearing sunglasses, back view of a cat, but the view is wearing sunglasses, etc... I actually tried using projective terms from drafting books, and didn't get great results. Nor anatomicals either. In short: natural language is not good enough and you need a DSL. If only the last 60 years of language research had warned us of this. Up next: English sentences are ambiguous and need context…

I mean... It kinda did. Sarcasm aside.

The trouble is all those darn uninitiated and trying to create a generalized oracle to map their inspecific ramblings to what they mean to free them of having to actually communicate properly...

Actually, funnily enough, this has cross section with philosophy in a way most programmers scoff at; but communication is frigging hard, and worse, detecting when someone is trying to get something across, but just needs a nudge in the right direction to be able to find the language to explain it is really damn hard.

I run into it every time I get a haircut. I have no idea how to speak their language, so it's always "Uh... A little off the top and rounded at the back, I guess?"

Re: DreamFusion: Text-to-3D using 2D Diffusion

#200

As someone who went to college for 3D animation in +* 1997* + AND DESIGNED the datacenter for luca' presidio complex.. where-by learning that Pixar was developed by steve jobs when lucas didnt think there was a future for computer animation... and so steve bought the death star from lucas... That became pixar... AI is going to fucking kill it - what will happen in the next decade will be ANYONE uploading a script to…

> AI is going to fucking kill it - what will happen in the next decade will be ANYONE uploading a script to an AI to make a full length movie...

Nah. These techniques will definitelly lower the barrier for making stuff (not just movies), but that has been case will all transformative technologies.

Before computers, if you wanted to shoot and edit a movie, it was a challenge. Now, you can shoot a movie with your pocket computer and edit it on the same device while shitting. Upload it to video sharing web and billions of people can watch it.

This class of technologies will enable creative people to make a lot of stuff, to iterate quickly. But don’t be naïve that everybody will do that. 99% of that will be trash, and that is fine. I like that it will enable individuals to bring their visions into this world without any need for collaboration. And when highly artistic individuals will begin to collaborate using these tools, that will be an awesome inflection point for art as we know it.

Post reply on HN