Live data from Hacker News

DreamFusion: Text-to-3D using 2D Diffusion

dreamfusion3d.github.io

91–100 of 208 posts

Re: DreamFusion: Text-to-3D using 2D Diffusion

#91
post #74
post #28

It's funny that the authors are 'anonymous' but they have access to Imagen so obviously it's by Google.

A large portion of the ML community (rightly) discredits Google papers because: - they rarely provide the data or code used so it's basically "i swear it works bro" research - what they achieve is usually through having the most pristine dataset on the planet and is often unusable by other researchers - other times they publish papers that are basically "we slightly modified this excellent open source paper, slapped…

A ton of big advances in AI that community has benefitted from have been from published google research

Re: DreamFusion: Text-to-3D using 2D Diffusion

#92

Earlier quoted context omitted.

The rolling pin is above the table but the shading is wrong because they don't render shadows.

Hi! Ajay here. Correct, our shading model doesn't compute intersections, since that's a bit challenging with a NeRF scene representation.

Super interesting work. Do you think that's a solvable problem and something you'll work on?

Re: DreamFusion: Text-to-3D using 2D Diffusion

#93
post #90
post #87

Earlier quoted context omitted.

Hello Stavros, I agree. When I look at the goals that base58 sought to achieve, (eliminating visually similar characters) I couldn't help but wonder why more characters were not eliminated. There is quite a bit of typeface androgyny when you consider case and face.

Yeah, I don't know why 1 was left in there, seems like a lost opportunity. Discarding l, I, 0, O, but then leaving 1? I wonder why.

I can only assume it was for a superstitious reason so that the original address prefixes could be a 1. This is the only sense I can make from it.

Re: DreamFusion: Text-to-3D using 2D Diffusion

#94

The most incredible thing here is that this demonstrates a level of 3D understanding that I didn't believe existed in 2D image models yet. All of the 3D information in the output was inferred from the training set, which is exclusively uncurated and unsorted 2D still images. No 3D models, no camera parameters, no depth maps. No information about picture content other than a text label (scraped from the web and often…

Co-author here - we were also surprised :) The breadth of knowledge of the visual world embedded in these 2D models and what they unlock is astounding.

Re: DreamFusion: Text-to-3D using 2D Diffusion

#95

The most incredible thing here is that this demonstrates a level of 3D understanding that I didn't believe existed in 2D image models yet. All of the 3D information in the output was inferred from the training set, which is exclusively uncurated and unsorted 2D still images. No 3D models, no camera parameters, no depth maps. No information about picture content other than a text label (scraped from the web and often…

So I wonder if unusual angles that normally do not get photographed will be distorted? For example, underneath a table looking up.

Yes, this is often a problem. We use view-dependent prompts (e.g. "cat wearing sunglasses, back view") but the pretrained 2D model often does not do a good job of interpreting non-canonical views and will put sunglasses on the back of the cats head (as well as the front).

Re: DreamFusion: Text-to-3D using 2D Diffusion

#98

The most incredible thing here is that this demonstrates a level of 3D understanding that I didn't believe existed in 2D image models yet. All of the 3D information in the output was inferred from the training set, which is exclusively uncurated and unsorted 2D still images. No 3D models, no camera parameters, no depth maps. No information about picture content other than a text label (scraped from the web and often…

So I wonder if unusual angles that normally do not get photographed will be distorted? For example, underneath a table looking up.

Some of the "mesh exports" used as examples on the page actually show this, to some extent. Look specifically at the plush corgi's belly and the weird geometry on the underside of the lemur's book, and to a lesser extent the underside of the bedsheet-ghost.

Re: DreamFusion: Text-to-3D using 2D Diffusion

#99
post #94

The most incredible thing here is that this demonstrates a level of 3D understanding that I didn't believe existed in 2D image models yet. All of the 3D information in the output was inferred from the training set, which is exclusively uncurated and unsorted 2D still images. No 3D models, no camera parameters, no depth maps. No information about picture content other than a text label (scraped from the web and often…

Co-author here - we were also surprised :) The breadth of knowledge of the visual world embedded in these 2D models and what they unlock is astounding.

Any word about how Nerf -> marching cubes works? I thought that was still an open problem. Is that another discovery in this research paper?

Re: DreamFusion: Text-to-3D using 2D Diffusion

#100

Earlier quoted context omitted.

They're AI generated, the singularity already happened but the machines are trying to ease us into it.

Scary fn thought! and I agree with you! And the OP comment its by the magnanimous/infamous AnigBrowl You need to start doing AI legal admin ( I dont have the terms, but you may - we need legal language to control how we deal with AI) and @dang - kill the gosh darn "posting too fast" thing Jiminey Crickets I have talked to you abt this so many times...

"You're posting too fast" is a limit that's manually applied to accounts that have a history of "posting too many low-quality comments too quickly and/or getting into flame wars". You can email dang (hn@ycombinator.com) if you think it was applied in error, but if it's been applied to you more than once... you probably have a continuing problem with proliferating flame wars or posting otherwise combative comments.
Post reply on HN