Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

11–20 of 373 posts

Re: No elephants: Breakthroughs in image generation

#11
post #2

That proper "no elephants" first image is hilarious. Another key point here is the generative AI's meme game is getting rather strong. Which isn't a small thing, humour is an advanced soft skill.

Having thousands of copies of that image in your training set isn’t a skill at all.

Re: No elephants: Breakthroughs in image generation

#12
> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent.

I have to disagree with the conclusion. This was an important discussion to have two to three years ago, then we had it online, and then we more or less agreed that it's unfair for artists to have their works sucked up with no recourse.

What the post should say is "we know that this is unfair to artists, but the tech companies are making too much money from them and we have no way to force them to change".

Re: No elephants: Breakthroughs in image generation

#13

Looking at the example where the coffee table is swapped, I notice every time the image is reprocessed it mutates, based on the previous iteration, and objects become more bizarre each time, like chinese whispers. * The weird-ass basket decoration on the table originally has some big chain links (maybe anchor chain, to keep the theme with the beach painting). By the third version, they're leathery and are merging wit…

The pictures on the wall change too.

Re: No elephants: Breakthroughs in image generation

#14
post #13

Looking at the example where the coffee table is swapped, I notice every time the image is reprocessed it mutates, based on the previous iteration, and objects become more bizarre each time, like chinese whispers. * The weird-ass basket decoration on the table originally has some big chain links (maybe anchor chain, to keep the theme with the beach painting). By the third version, they're leathery and are merging wit…

The pictures on the wall change too.

Yes, first a still life and something impressionist, then a blob and a blob, then a smear and a smear. And what about the reflections and transparency of the glass table top? It gets very indistinct. Keep working at the same image and it looks like you'll end up with some Deep Dream weirdness.

I think the fireplace might be turning into some tiny stairs leading down. :)

Re: No elephants: Breakthroughs in image generation

#15

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

> This was an important discussion to have two to three years ago, then we had it online, and then we more or less agreed that it's unfair for artists to have their works sucked up with no recourse.

Speak for yourself, there was no consensus online. There are plenty of us that think that dramatically expanding the power of copyright would be a huge mistake that would primarily benefit larger companies and do little to protect or fund small artists.

Re: No elephants: Breakthroughs in image generation

#16
post #5
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

I've never used a stock photo site before, so I suppose it's no surprise I have no real use for "generate any image on demand".

Their main application appears to be taking blog posts and internal memos and making them three times longer and use ten times the bandwidth to convey no more information. So exactly the application AI is ‘good’ at.

Re: No elephants: Breakthroughs in image generation

#17
post #13

Earlier quoted context omitted.

The pictures on the wall change too.

Yes, first a still life and something impressionist, then a blob and a blob, then a smear and a smear. And what about the reflections and transparency of the glass table top? It gets very indistinct. Keep working at the same image and it looks like you'll end up with some Deep Dream weirdness. I think the fireplace might be turning into some tiny stairs leading down. :)

Only sailors know how to leave.

Re: No elephants: Breakthroughs in image generation

#18
post #4

I had a reasonable intuition for how the "old" method works, but I still don't grok this new approach. "in multimodal image generation, images are created in the same way that LLMs create text, a token at a time" Is there some way to visualise these "image tokens", in the same way I can view tokenized text?

Imagine you cut the image into 32x32 pixel blocks. And then for each block, you can chose 1 out of 128,000 variations. And then a post-processing step smoothes out the borders between blocks and adjusts small details. That's basically how a transformer image generation model works. As such, the process is remarkably similar to old fixed-font ASCII art. It's just that modern AIs have a larger alphabet and, thus, more…

I don't get how this would produce consistent images. In the article, the text could be on a grid, but the window and doorway and sofa don't seem to be grid-aligned. (Or maybe the text is overlaid?)

Re: No elephants: Breakthroughs in image generation

#19
post #15

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

> This was an important discussion to have two to three years ago, then we had it online, and then we more or less agreed that it's unfair for artists to have their works sucked up with no recourse. Speak for yourself, there was no consensus online. There are plenty of us that think that dramatically expanding the power of copyright would be a huge mistake that would primarily benefit larger companies and do little t…

>There are plenty of us that think that dramatically expanding the power of copyright would be a huge mistake that would primarily benefit larger companies and do little to protect or fund small artists.

The status quo also primarily benefits larger companies, and does little (exactly nothing, if we're being earnest) to protect or fund small artists.

It's reasonable to hold both opinions that: 1) artists aren't being compensated, even though their work is being used by these tools, and 2) massive expansion of copyright isn't the appropriate response to 1).

Post reply on HN