Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

41–50 of 373 posts

Re: No elephants: Breakthroughs in image generation

#41

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

>> it's unfair for artists to have their works sucked up

What framework can we use to decide if something is fair or not?

Style is not something that should be copyrighted. I can pain in the style of X painter, I can write in the style of Y writer, I can compose music in the style of Z composer.

Everything has a style. Dressing yourself has a style. Speaking has a style. Even writing mathematical proofs can have a style.

Copying another person's style might reflect poor judgement, bad taste and lack of originality but it shouldn't be illegal.

And anyone in the business of art should have much more than a style. He should have original ideas, a vision a way to tell stories, a way to make people ask themselves questions.

A style is merely a tool. If all someone has is a style, then good luck!

Re: No elephants: Breakthroughs in image generation

#42
post #4

I had a reasonable intuition for how the "old" method works, but I still don't grok this new approach. "in multimodal image generation, images are created in the same way that LLMs create text, a token at a time" Is there some way to visualise these "image tokens", in the same way I can view tokenized text?

I haven't see any details on how OpenAI's model works, but the tokens it generates aren't directly translated into pixels - those tokens are probably fed into a diffusion process which generates the actual image.. The tokens are the latent space or conditioning for the actual image generation process.

> I haven't see any details on how OpenAI's model works

Exactly. People just confidently make things up. There are many possible ways, and without details, "native generation" is just a marketing buzzword without clear definition. It's a proprietary system, there is no code release, there is no publication. We simply don't know how exactly it's done.

Re: No elephants: Breakthroughs in image generation

#43
post #38
post #34

It's interesting to hear people side with the artists when in previous discussions on this forum I've gotten significant approval/agreement arguing that copyright is far too long. As I've argued in the past, I think copyright should last maybe five years: in this modern era, monetizing your work doesn't (usually) have to take more than a short time. I'd happily concede to some sort of renewal process to extend that p…

I find it unlikely that someone who was willing to pay an artist for a charcoal sketch would be satisfied with an AI alternative. You don't just buy art for the aethstetic, you buy it for a lot of reasons and AI doesn't give any of the same satisfaction.

I'm all for paying artists for their work. Unfortunately, same as tattoo artists, some just heavily overcharge for mediocre results (been tattooing myself AND I know a few things about art). Like, sorry, but if you want to earn money doing art, please be good at it...

Re: No elephants: Breakthroughs in image generation

#44
post #34

It's interesting to hear people side with the artists when in previous discussions on this forum I've gotten significant approval/agreement arguing that copyright is far too long. As I've argued in the past, I think copyright should last maybe five years: in this modern era, monetizing your work doesn't (usually) have to take more than a short time. I'd happily concede to some sort of renewal process to extend that p…

Being someone who has paid a lot of attention to Ghibli, I wouldn’t say their style was well established 35-40 years ago… There is considerable evolution and refinement to their style from Naushika, to later works, both in the artistic style and the philosophical content it presents.

I think allowing it to be fair game would have destroyed something quite beautiful that I’ve watched evolve across 40 years and which I was hoping to see the natural conclusion of without him being bothered by the AI-fication of his work.

Re: No elephants: Breakthroughs in image generation

#45

Are there any local models that use this new approach to generating images?

GPT-4o is the only model that seems to work well in the text-image joint space to this degree, even Gemini Flash 2.0 with native image support is not nearly as good so it will probably be a while for a good open source alternative to pop up (a while in the context of AI development).

Re: No elephants: Breakthroughs in image generation

#46
post #33

> Is it okay to reproduce the hard-won style of other artists using AI? Who owns the resulting art? Who profits from it? Which artists are in the training data for AI, and what is the legal and ethical status of using copyrighted work for training? These were important questions before multimodal AI, but now developing answers to them is increasingly urgent. I have to disagree with the conclusion. This was an importa…

I don't think there's consensus around that idea. Lots of people (myself included) feel that copyright is already vastly overreaching, and that AI represents forward progress for the proliferation of art in society (its crap today, but digital cameras were crap in 2007 and look where they are now). Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet. I went…

>Its also not clear for example that Studio Ghibli lost by having their art style plastered all over the internet.

Maybe Studio Ghibli is much more than merely a style. Maybe people aren't looking at their production just for the style.

Most people dislike wearing fake clothes and the dislike wearing fake watches or fake jewelry. Because it isn't just about the style.

Re: No elephants: Breakthroughs in image generation

#47
post #43
post #38

Earlier quoted context omitted.

I find it unlikely that someone who was willing to pay an artist for a charcoal sketch would be satisfied with an AI alternative. You don't just buy art for the aethstetic, you buy it for a lot of reasons and AI doesn't give any of the same satisfaction.

I'm all for paying artists for their work. Unfortunately, same as tattoo artists, some just heavily overcharge for mediocre results (been tattooing myself AND I know a few things about art). Like, sorry, but if you want to earn money doing art, please be good at it...

> Like, sorry, but if you want to earn money doing art, please be good at it...

By definition, almost half of all $ARTISTS are worse than the median. Should that half not get paid for their time?

Re: No elephants: Breakthroughs in image generation

#48
> The question isn't whether these tools will change visual media, but whether we'll be thoughtful enough to shape that change intentionally.

Unfortunately I think the answer to this question is a resounding “no”.

The time for thoughtful shaping was a few years ago. It feels like we’re hurtling toward a future where instead we’ll be left picking up the pieces and assessing the damage.

These tools are impressive and will undoubtedly unlock new possibilities for existing artists and for people who are otherwise unable to create art.

But I think it’s going to be a rough ride, and whatever new equilibrium we reach will be the result of much turmoil.

Employment for artists won’t disappear, but certain segments of the market will just use AI because it’s faster, cheaper, and doesn’t require time consuming iterations and communication of vision. The results will be “good enough” for many.

I say this as someone who has found these tools incredibly helpful for thinking. I have aphantasia, and my ability to visualize via AI is pretty remarkable. But I can’t bring myself to actually publish these visualizations. A growing number of blogs and YouTube channels don’t share these qualms and every time I encounter them in the wild I feel an “ick”. It’ll be interesting to see if more people develop this feeling.

Re: No elephants: Breakthroughs in image generation

#50

Earlier quoted context omitted.

I haven't see any details on how OpenAI's model works, but the tokens it generates aren't directly translated into pixels - those tokens are probably fed into a diffusion process which generates the actual image.. The tokens are the latent space or conditioning for the actual image generation process.

> I haven't see any details on how OpenAI's model works Exactly. People just confidently make things up. There are many possible ways, and without details, "native generation" is just a marketing buzzword without clear definition. It's a proprietary system, there is no code release, there is no publication. We simply don't know how exactly it's done.

Open AI have both said it's native image generation and autoregressive. It has the signs of it too.

It's probably an implementation of VAR (https://arxiv.org/abs/2404.02905) - autoregressive image generation with a small twist. Rather than predict every token at the target resolution directly, start with predicting it at a small resolution, cranking it higher and higher until the desired resolution.

Post reply on HN