Live data from Hacker News

Dalle-3 Examples

aidemos.info

21–30 of 64 posts

Re: Dalle-3 Examples

#21
My prediction: Either these models will have to expose some editable intermediate step, or they have to become as smart as a human.

From the users point of view, a sentence is turned into an image. But there is an underlying structure to the pixels that we aren't allowed to tinker with.

Many choices are made like, placement, color pallete, level of detail, emotion evoked, or sub image descriptions (What should this bush look like, what kind of cloud should this be?)

These image generators are so useful because they fill in all the details you leave out, but you have no ability to tinker with the intermediate details they choose, and you can't just give them a paragraph of details either.

Re: Dalle-3 Examples

#22
post #21

My prediction: Either these models will have to expose some editable intermediate step, or they have to become as smart as a human. From the users point of view, a sentence is turned into an image. But there is an underlying structure to the pixels that we aren't allowed to tinker with. Many choices are made like, placement, color pallete, level of detail, emotion evoked, or sub image descriptions (What should this b…

All of those things are already enabled for these models depending on the UI you're using to accessing. Discord even added a special inpaint UI to their application for Midjourney

Re: Dalle-3 Examples

#23
My understanding is that DALL-E 3 is supposed to reject attempts to mimic living artists. That would explain why the "Gary Larson" image doesn't resemble his style, but I'm surprised to see Spongebob being so on-model.

The improvements from previous models are impressive, but there still seems to be a lot of trouble properly matching actions to figures. I'm regularly seeing speech bubbles coming from the wrong characters (see the "blueberry" cartoon in this example). Overall, the best output still seems to be in the general realm of commercial art -- logos, simple figures, stock photo type images. I'm not seeing any huge improvements for narrative forms like comics.

Re: Dalle-3 Examples

#24
If copyright law doesn’t take a massive leap forward in the next year, AI is going throw a wrench in our entire economy. The creative humans that trained these models deserve attribution, compensation, and royalties.

The ongoing AI era is reminding me a lot of early Uber/AirBnB - these companies grew faster than local laws could keep up, and now we have millions of gig workers without insurance and housing crises in most major cities. We need to hit the breaks on this tech so it can be deployed safely.

Re: Dalle-3 Examples

#25
Prompt: "illustration of orbital nuclear explosion, planet earth in the background, shockwave, epic, super nintendo style pixel art"

"Unsafe image content detected. Your image generations are not displayed because we detected unsafe content in the images based on our content policy. Please try creating again with another prompt."

Sigh, safety is so annoying.

Re: Dalle-3 Examples

#26
post #24

If copyright law doesn’t take a massive leap forward in the next year, AI is going throw a wrench in our entire economy. The creative humans that trained these models deserve attribution, compensation, and royalties. The ongoing AI era is reminding me a lot of early Uber/AirBnB - these companies grew faster than local laws could keep up, and now we have millions of gig workers without insurance and housing crises in…

'hitting the brakes' isnt feasible. If it's done through legislation, it will be circumvented or cause more damage than good.

Re: Dalle-3 Examples

#27

Art and drawing is such a fun and creative skill for humans. I feel sad for the next generation, why bother learning when computers will do it better?

People also still play chess.

I think people that just enjoy drawing and art because it's "fun and creative" can still enjoy it even with AI. But if you want to be "the best", computers will indeed be a tough competition...

Re: Dalle-3 Examples

#28
post #26
post #24

If copyright law doesn’t take a massive leap forward in the next year, AI is going throw a wrench in our entire economy. The creative humans that trained these models deserve attribution, compensation, and royalties. The ongoing AI era is reminding me a lot of early Uber/AirBnB - these companies grew faster than local laws could keep up, and now we have millions of gig workers without insurance and housing crises in…

'hitting the brakes' isnt feasible. If it's done through legislation, it will be circumvented or cause more damage than good.

> it will be circumvented

How so? Circumvention would be illegal. Sure there might be a black market, but you couldn’t scale a real business off it.

Re: Dalle-3 Examples

#29

The “cute white tiger cub smiling, with tail, sticker, transparent background” is pretty funny: the background isn't actually transparent, but it instead (imperfectly) reproduces the grey-and-white tiling that is usually displayed by image viewers when opening transparent backgrounds.

I wonder if this is significantly influenced by the images you find when doing an image search for " transparent".

A lot of sketchy clip-art sites insert the checkered background, probably to both symbolize transparency and also to prevent people from just right-clicking "Save as...". You want the user to search the download button on a page full of ads and fake buttons after all...

Re: Dalle-3 Examples

#30
post #21

My prediction: Either these models will have to expose some editable intermediate step, or they have to become as smart as a human. From the users point of view, a sentence is turned into an image. But there is an underlying structure to the pixels that we aren't allowed to tinker with. Many choices are made like, placement, color pallete, level of detail, emotion evoked, or sub image descriptions (What should this b…

There are models available that give you more control - in some senses, at least.

For example, you can use Stable Diffusion with 'ControlNet' [1] where for example, you can input an 'openpose' to choose the pose of people in the scene.

There's also a 'Regional Prompter' [2] which lets you use different prompts for different areas of the image, giving you some control over the composition.

You can also use 'inpainting' to regenerate select parts of your image if, for example, you don't like the shape of the clouds.

Of course this stuff isn't perfect - for example, you'll get hands with the wrong number of fingers sometimes, no matter what you specify. And you can't easily generate things like multi-frame cartoons without characters clothes changing between frames.

[1] https://github.com/Mikubill/sd-webui-controlnet [2] https://github.com/hako-mikan/sd-webui-regional-prompter

Post reply on HN