Live data from Hacker News

No elephants: Breakthroughs in image generation

oneusefulthing.org

341–350 of 373 posts

Re: No elephants: Breakthroughs in image generation

#341

Earlier quoted context omitted.

> What's going to happen when someone can make a competing movie in their style with just a prompt? Nothing? Just like how if some studio today invests millions of man-hours and does a competing movie in Studio Ghibli's aesthetic (but not including any Studio Ghibli's characters, branding, etc. - basically, not the copyrightable or trademarkable stuff) nothing out of ordinary is going to happen. I mean, artistic styl…

You are missing the point entirely. If you can make a movie with just a prompt, who is going to invest the money creating something like a Ghibli movie just to have it ripped off? Instead people will just rip off what has already been done and everything just stagnates.

How is "movies and other great works of art that used to cost tens of millions of dollars to make now cost tens of dollars to make" a bad thing?

It means art can get more ambitious. Ghibli made their mark, and made their money. Now it's time for the next generation to have a turn.

Re: No elephants: Breakthroughs in image generation

#342

Earlier quoted context omitted.

I've used Midjourney and chatGPT. Midjourney is better for rapid iteration, cycling through options faster, and to a large extent getting "weirder". It's easier to tweak using parameters. ChatGPT is far, far superior (especially now) when you want something more specific that you've already imagined. But it's slower, and unlike Midjourney you don't get four versions to choose to build and iterate on, you get a single…

> four versions to choose to build and iterate on How does this work? How do you ask a model to produce four different variations? Or do they have four different models run the same inference?

There's some noise in the process, so you won't get the same results if you ask the same model the same prompt 4 times. Most of the services that do this just ask the same model 4 times with a different random seed, as far as I can tell.

Re: No elephants: Breakthroughs in image generation

#343

Earlier quoted context omitted.

I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results, which is good enough for the presentations I hold, online pet projects (created a squirrel-themed online math-learning game for which I previously would have needed a designer to create squirrel highschool themed imagery)…

>I love using LLMs to generate pictures. I'd call myself rather creative, but absolutely useless in any artistic craft. Now, I can just describe any image I can imagine and get 90% accurate results May I ask what you use? I'm not yet even a paid subscriber to any of the models, because my company offer a corporate internal subscription chatbot and code integration that works well enough for what I've been doing so fa…

The new ChatGPT image generation is insane. It's available on the free tier, just strongly rate limited.

Re: No elephants: Breakthroughs in image generation

#344

Earlier quoted context omitted.

I've used Midjourney and chatGPT. Midjourney is better for rapid iteration, cycling through options faster, and to a large extent getting "weirder". It's easier to tweak using parameters. ChatGPT is far, far superior (especially now) when you want something more specific that you've already imagined. But it's slower, and unlike Midjourney you don't get four versions to choose to build and iterate on, you get a single…

> four versions to choose to build and iterate on How does this work? How do you ask a model to produce four different variations? Or do they have four different models run the same inference?

You don't ask to run four models, it's just that every single prompt you give it will return four models. You can then choose to have one of them iterated on, or upscaled.

If you're still not sure let me know and I'll show a screenshot.

Re: No elephants: Breakthroughs in image generation

#345
post #246

Earlier quoted context omitted.

>But I think it’s going to be a rough ride, and whatever new equilibrium we reach will be the result of much turmoil. Honestly visual media just seems to be the start. In the past two years we've seen about as much robotics progress as the last 20. If this momentum keeps up then we're not just talking about artists that are going to have issues.

Honestly, I'm pretty encouraged by all of the projects and efforts within legislation and organizations regarding clear lines being drawn - i.e., through watermarking to clearly label whether something is AI-generated or not - as well as the efforts by industries for livelihoods to be protected, specifically in the creative space, where human intentionality and feeling are still of the utmost essentiality. We've seen…

While I know there have been plenty of scathing essays, backlash among various communities, etc. do you have some concrete examples of the clear lines being drawn and legislation that gives you this optimism?

Maybe the progress you’re describing has escaped me because of the sheer speed this is all unfolding, but it feels like all I’ve heard is lots of noise, while AI companies continue to hammer hosted resources across the Internet to build their next model, the US government continues to claim they’ll use AI to solve problems of waste and fraud, companies like Shopify claim they won’t hire anyone unless it can be proven that AI cannot do the job, and an increasing % of the content I encounter is AI slop.

Maybe this is all necessary for a proper backlash to form, and I definitely want to become more aware of the positives anywhere I can find them. I’m not an AI doomer, but haven’t yet found the optimism you describe.

Re: No elephants: Breakthroughs in image generation

#346
post #3

This is a before/after moment for image generation. A simple example is the background images on a ton of (mediocre) music youtube channels. They almost all use AI generated images that are full of nonsense the closer you look. Jazz channels will feature coffee shops with garbled text on the menu and furniture blending together. I bet all of that disappears over the next few months. On another note, and perhaps other…

> But now that they're here, I just sort of poke at it for a minute and carry on with my day.

Well, that's because they suck, despite all the hype.

They have a use in a professional context, i.e., as replacement for older models and algorithms like BERT or TF/IDF.

But as assistants they're only good as a novelty gag.

Re: No elephants: Breakthroughs in image generation

#347
post #340

I'm more interested in the technical details than the publicity. Pretty much anyone these days can learn what a diffusion model is, how they're implemented, what the control flow is. What about this new multimodal LLM? They have no problems with text, they generate images using tokens, but how exactly? There's no open-source implementations that I know of, and I'm struggling to find details.

This video is very good. https://youtu.be/EzDsrEvdgNQ?si=EWp3U1GMkwg1bMQQ

One thing i'd add is that generating the tokens at the target resolution from the start is no longer the only approach to autoregressive image generation.

Rather than predicting each patch at the target resolution right away, it starts with the image (as patches) at a very small resolution and increasingly scales up. Paper here - https://arxiv.org/abs/2404.02905

Re: No elephants: Breakthroughs in image generation

#348

Earlier quoted context omitted.

I searched Shutterstock for squirrels doing math. Here's a squirrel doing math: https://www.shutterstock.com/image-vector/pensive-squirrel-d... Yes, it's obvious that if your use case is obscure enough, or you need a ton of unique images, they won't work, which is why I said "largely a solved problem".

But image generation these days is simply image search anyway. Wether or not the image existed before is almost irrelevant.

I've never searched for an image from a stock vendor the way one would prompt a generative model. The stock vendor's metadata/keyword about its images were never that in-depth.

Also, you're implying that a generative system is so fast that it could create so many variations of your prompt to fill in the search results page in an acceptable time. That's a joke

Re: No elephants: Breakthroughs in image generation

#349
post #341

Earlier quoted context omitted.

You are missing the point entirely. If you can make a movie with just a prompt, who is going to invest the money creating something like a Ghibli movie just to have it ripped off? Instead people will just rip off what has already been done and everything just stagnates.

How is "movies and other great works of art that used to cost tens of millions of dollars to make now cost tens of dollars to make" a bad thing? It means art can get more ambitious. Ghibli made their mark, and made their money. Now it's time for the next generation to have a turn.

> How is "movies and other great works of art that used to cost tens of millions of dollars to make now cost tens of dollars to make" a bad thing?

It's bad because you will never get an original visual style from now on. Everything will be copy-paste of existing styles, forever.

Re: No elephants: Breakthroughs in image generation

#350

Earlier quoted context omitted.

Right the difference is that it’s a large company looking at it then copying it and reselling it without credit, which basically everyone would understand as bad without the indirection of a model. Edit: the key words here are “company” and “reselling”

But it’s not copying and reselling — it’s imitation. Copying is controlled by copyrights. And imitation isn’t controlled by anything. As for a company: a company is just a group of people acting together.

Try "imitating" some Mario Brothers in a commercial context and see how that goes. Good luck.
Post reply on HN