Live data from Hacker News

Dall-E 2

openai.com

341–350 of 511 posts

Re: Dall-E 2

#341

Earlier quoted context omitted.

This thing can't do 3D models. There are some 3D image generation techniques, but they aren't based on polygonal modelings, so 3D artists are safe for now

You could train a model on texture image data though, no? Or what about even generating images you could then photogrammetry into models?

Yeah, there's a lot of 2D assets that this model would be great for (textures, materials, *maps, etc) that would definitely improve the asset-building process for game devs. I've already used VQGAN+CLIP for some low-res skill and item icons in hobby games and it seems things are only improving from here.

I wouldn't be surprised to see a comparable version for 3D models in the next year or two, though. Even if the current architecture doesn't lend itself to 3D structures (I don't know), there's a lot of parallel work being done right now (esp. by Google) for encoding 3D data in new/efficient ways, translating specialized 2D images into 3D models, and more.

Re: Dall-E 2

#342

Earlier quoted context omitted.

Not quite, it looks like this: - Provide an existing image - Provide a text prompt ("flamingo") - Select from X variations the new image that looks best to you - It does the equivalent of a google image search on your "flamingo" prompt - It picks the most blend-able ones as a basis to a new synthetic flamingo - It superimposes the result on your image Very cool don't get me wrong. Now I want to tweak this new floatin…

That's not how this works. There is no 'search' step, there is no 'superimposing' step. It's not really possible to explain what the AI is doing using these concepts. If you pay attention to all the corgi examples, the sofa texture changes in each of them, and it synthesizes shadows in the right orientation - that's what it's trained to do. The first one actually does give you the impression of weight. And if you loo…

How do you propose we talk about what it is doing if not by using the terminology from the human editing process it is replacing? I'm struggling to express things.

My issue is that it appears to not be possible to explain what the AI is doing at all. If you could, you'd be able to actually control the output. And talking about how the model is trained is interesting but not an answer.

Of course there is a superimposing step, that just means it adds its layer on top of the photo you provide. That's all it means and that's literally what it is doing, that's all I tried to say, heh.

> If you pay attention to all the corgi examples, the sofa texture changes in each of them

Yes, exactly!

> This is far from a 'cool trick' and many of those images would take hours for a human to reproduce

OK, fair enough. I'll try to be more clear:

It is very cool and not a trick and the results are fantastic if you got out exactly what you wanted. Amazing time saver. And if not? Right now this is totally hit or miss.

It would also take hours for a human to reproduce a Vermeer and this no doubt has those in its training set and would style-transfer unto a corgi instantly. Certainly faster than Vermeer himself could do it.

But Vermeer could explain how he came up with the style, his techniques, choices, 'etc.

It reads like the advance here is that it will usually synthesize something that looks great but not always the thing that you want. With no recourse.

Re: Dall-E 2

#343

Earlier quoted context omitted.

This is exactly what they demo - they lock a scene and add a flamingo in three different locations. In another one they lock the scene and add a corgi.

Not quite, it looks like this: - Provide an existing image - Provide a text prompt ("flamingo") - Select from X variations the new image that looks best to you - It does the equivalent of a google image search on your "flamingo" prompt - It picks the most blend-able ones as a basis to a new synthetic flamingo - It superimposes the result on your image Very cool don't get me wrong. Now I want to tweak this new floatin…

Are your affirmations based on the content of the paper ?

Re: Dall-E 2

#344
post #43

Something about this makes me nauseous. Perhaps is the fact that soon the market value for creatives is going to fall to a hair about zero for all but the most famous. We will be all the poorer for it when 95% of images you see are AI generated. There will be niches of course but in a few short years it'll be over for a huge swathe of creative professionals who are already struggling. Some of the images also hit me w…

I really don't agree. When I work with a creative I'm not working with them because of their content generation skills. I'm working with them because of their taste and curation ability that results in the end product. The nature of creative work will certainly change, creatives will adopt tools such as Dall-E 2. In certain narrow cases they might be replaced, such as if you are asking a creative to generate a very s…

>I'm working with them because of their taste and curation ability that results in the end product. ... The nature of creative work will certainly change, creatives will adopt tools such as Dall-E 2.

Furthermore, tools like Dall-E seem like they'll lower the barrier of entry for more people to get into art, resulting in more artists, not fewer. Increased competition for the same dollar amounts might make artists, on average, "poorer" (when averaged across an increased number of artists), but this seems like the end-result of any new tool that empowers more artists to more easily make "good" work, not just AI-generated tools.

I'm excited for both 1) more art in the world and 2) in some cases, artists making "even better" art (by combining their existing experience + new tools).

Re: Dall-E 2

#345

I can see how this has the potential to disrupt the games industry. If you work on a AAA title, there is a small army of artists making 19 different types of leather armor. Or 87 images of car hubcaps. Using something like this could really help automate or at least kickstart the more mundane parts of content creation. (At least when you are using high resolution, true color imagery.)

This thing can't do 3D models. There are some 3D image generation techniques, but they aren't based on polygonal modelings, so 3D artists are safe for now

Colab notebooks can do mesh models using this method, I'm certain OpenAI isn't far away

Re: Dall-E 2

#346

Earlier quoted context omitted.

https://www.eleuther.ai (text, not images, but free as in freedom)

Katherine Crawson is @ Eletheur & IMHO is indisputably most responsible for the advances in text=>image generation. Dall-E 2 is Dall-E and her insight to use diffusion, the intermediate proof of concept of diffusion + Dall-E is GLIDE. https://twitter.com/RiversHaveWings & https://github.com/crowsonkb

Diffusion models existed long before this announcement. I have no idea who this person is, but they did not invent this idea.

Edit: Diffusion models guided by CLIP*

Re: Dall-E 2

#347

>We’ve limited the ability for DALL·E 2 to generate ... adult images. I think that using something like this for porn could potentially offer the biggest benefit to society. So much has been said about how this industry exploits young and vulnerable models. Cheap autogenerated images (and in the future videos) would pretty much remove the demand for human models and eliminate the related suffering, no? EDIT: typo

Depends whether you think models should be able to generate cp. It's almost impossible to even give an affirmative answer to that question without making yourself a target. And as much as I err on the side of creator freedom, I find myself shying away from saying yes without qualifications. And if you don't allow cp, then by definition you require some censoring. At that point it's just a matter of where you censor,…

Of course some level of censorship is needed, otherwise it can be used to produce porn involving real people without their consent (eg celebs)

Re: Dall-E 2

#348

Earlier quoted context omitted.

Yes. Translating business requirements, customer context, engineering constraints, etc. into usable, practical, functional code, and then maintaining that code and extending it is so far beyond the horizon, that many other skillsets will replaced before programming is. After all, at that point, the AI itself, if it's so smart, should be able to improve itself indefinitely. In which case we're fucked. Programming will…

The tail end of programming will be the last thing to be replaced, maybe. I don’t see why CRUD apps get to hide under the umbrella of programming ultra-advanced AI.

Let me know when you can speak English to a computer and have it generate CRUD code that satisfies all engineering and design constraints. The AI will need to be dynamic enough to understand nuance, missing gaps in the requirements spec, have context on the application being built, able to suggest improvements on product design, know how to make changes through the same conversational interface, etc.

Accomplishing that is achieving general AI.

In the meantime, there are plenty of boilerplate ORMs and simplistic API template tools that make production of bog standard CRUD apps dead simple. Of course, they all have their drawbacks and trade-offs, and aren't always suitable. But I don't see the amount of software engineering work reducing as a result of these no-code, low-code tools, do you?

Re: Dall-E 2

#349
post #338

Earlier quoted context omitted.

Depends whether you think models should be able to generate cp. It's almost impossible to even give an affirmative answer to that question without making yourself a target. And as much as I err on the side of creator freedom, I find myself shying away from saying yes without qualifications. And if you don't allow cp, then by definition you require some censoring. At that point it's just a matter of where you censor,…

I don't think it's necessarily certain villainy for those who fight that fight as long as they are fighting it correctly. There's a huge case to be made that flooding the darknet with AI generated CP reduces the revictimization of those in authentic CP images, and would cut down on the motivating factors to produce authentic CP (for which original production is often a requirement to join CP distribution rings). As w…

How do you suppose your CP generator will be trained without using authentic CP images? Not only will that require revictimization but you’ll also be downloading CP to train the model.

Re: Dall-E 2

#350

Could somebody build this for SVG icons? I’d invest in it.

What do you want?

A library like https://icons8.com/icons where you can just tell it what icon you want and the style (e.g. Material, outline, solid, iOS). It would do it’s thing and spit it out.
Post reply on HN