Live data from Hacker News

SDXL Turbo: A Real-Time Text-to-Image Generation Model

stability.ai

131–140 of 157 posts

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#131

Earlier quoted context omitted.

Isn't literally every imagegen AI that's not DALL-E or Midjourney based on Stable Diffusion?

Are we sure that those arent based on stable diffusion? No code black box, and we get to tease the closed source companies for wrapping FOSS stuff. Midjourny I'm most convinced is just a SD with a fine-tuned model. That would explain why everything looks like pixar and can't follow the prompt.

Given that Midjourney predates StableDiffusion, that seems unlikely, though it is possible they threw away all their hard work to create their model to use one that's available to other people for free and then charge money for it.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#132
post #112

Earlier quoted context omitted.

> The US breaking Google up also won't end the monopoly money. It's also weird because modern economics with tech has created a space that creates a lot of natural monopolies. Momentum is very powerful and it's the reason silicon valley companies will run at a loss for years creating a userbase. Trick is to keep them (or sell before buyer starts charging). There's what, 2 map companies and only one of them is in high…

Seems like a great place to mention the MapQuest API. I tested it against 4 other services(inc. Google's) on 200 manually verified geocoding/reverse-geocoding tasks. Google managed a 92% whereas MapQuest scored 99%. If you are doing geocoding or reverse-geocoding MapQuest outperforms Google Maps quite magnificently(YMMV). The cost is also lower and there are plans where you can keep the data. The funny thing being th…

What were the others? Google Maps, MapBox, OSM are the three that come to my mind.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#133

Earlier quoted context omitted.

Yes - https://www.reuters.com/legal/ai-generated-art-cannot-receiv...

Actually, no. This is just another example of a headline leaving out important details of the actual case. In this case the plaintiff actually named the AI as the producer, not themselves. From the case: "the sole issue of whether a work generated entirely by an artificial system absent human involvement" This leaves a lot of wiggle room for AI created art with some form of human involvement. I.E. import the generate…

it doesn't touch trademarks either. If I generate a recognizable picture of a trademarked icon, I still can't use it for my own commercial purposes, even if it's not covered under copyright.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#134
post #90

Earlier quoted context omitted.

These models can't exist without the training sets. Their value is entirely derived from existing data. The ml architecture does not matter at all. Sure, throw enough compute and data at a problem, do a little parallelization, and you can extract plenty of patterns. Does that mean the ml engineers understand art? Or are they just using glorified brute force to alienate people who actually make things from their labor…

Do humans understand art? Does it matter? How do humans learn to create art? By looking at other peoples art, for the most part. Those other artists are not compensated for this either, nor do they need to be credited. If you exactly copy another persons art style, that may be frowned upon by some, but otherwise it is of no consequence, unless you claim the work is actually made by that artist . If you believe that h…

You're the millionth person to make the argument that we should treat AI learning the same as human learning. It's still a bad argument, but I'm getting very tired of explaining why treating a human the same as a computer program sucks ass.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#135
post #134

Earlier quoted context omitted.

Do humans understand art? Does it matter? How do humans learn to create art? By looking at other peoples art, for the most part. Those other artists are not compensated for this either, nor do they need to be credited. If you exactly copy another persons art style, that may be frowned upon by some, but otherwise it is of no consequence, unless you claim the work is actually made by that artist . If you believe that h…

You're the millionth person to make the argument that we should treat AI learning the same as human learning. It's still a bad argument, but I'm getting very tired of explaining why treating a human the same as a computer program sucks ass.

You're the millionth person to make the argument that AI art sucks. That didn't stop me from typing up a reply for you. Why don't you link one of your previous replies?

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#136

Earlier quoted context omitted.

> The US breaking Google up also won't end the monopoly money. It's also weird because modern economics with tech has created a space that creates a lot of natural monopolies. Momentum is very powerful and it's the reason silicon valley companies will run at a loss for years creating a userbase. Trick is to keep them (or sell before buyer starts charging). There's what, 2 map companies and only one of them is in high…

Natural monopoly means that the market only has enough demand for one supplier. Canonical examples would include highways, railways, local telephone networks, and residential Internet access providers. Most of the big tech companies we love to hate aren't natural monopolies. Google's anticompetitive moat is primarily made of inertia: they were the first to market with a halfway functional search index. Defaults are v…

> Natural monopoly means that the market only has enough demand for one supplier.

That's not true. Natural monopolies also form due to network effects. You can follow the story of Bell Labs for an early example, which I suspect you're aware of. The reason for a monopoly wasn't for lack of demand it was about tragedy of the commons.

In our modern tech economics we have similar tragedy of the commons but a bit more abstracted. The thing is most products are a result of their userbase, not the product itself. Look at HN. Or look at Reddit or YouTube. It may look like circular logic (because it is a feedback loop), but the fact that everyone is publishing on YouTube makes YouTube bigger and more useful to the consumer which causes more people to publish on that platform because there are more users. It then makes it very difficult to compete because what can you do? You can make a platform that is 100x better than YouTube in respects to serving videos, search, pay to creators, and so on, but you won't win because you have no users and you won't have users without creators who aren't going to produce on your platform because there are no users. In other words, there's a first mover disadvantage. You're an early user and the site sucks because it has no content but you're a true believer. You're an early creator and your pay sucks because there are so few viewers (even if your pay per viewer is 100x your pay per video/time spent is going to be 10000x less because there are 1000000x fewer users). Look at PeerTube, Nebula, or FloatPlane. They are all doing fine but nowhere near as successful as YouTube which everyone hates on (and for good reason). Hell, when YouTube started trying to compete with Twitch they had to REALLY incentivize early creators to move with very lucrative deals because they were not buying the creator, they were buying their userbase. It should be a clear signal that there's an issue if a giant like Google has a hard time competing with Amazon.

For a highly competitive market you need a low barrier to entry so that you can disrupt. There are thousands of examples where a technology/product that is superior in every way (e.g. price and utility) but are not the market winners because network effects exist. Even things like BetaMax vs VHS is a story of network effects (I wouldn't say BetaMax dominated VHS, but not important), because what mattered was what you could get at a store or share (via your neighbor or local rental).

And I'm glad you mention Firefox, because it's a good example of stickiness. I've tried to convert many friends who groan and moan about how hard it is and make up excuses like bookmarks and literally showing them that on startup it'll export for you they just create a new excuse or say UI/UX is trash because the settings button is 3 horizontal lines instead of 3 vertical dots so they can't find it despite being in the same place or tabs are not as curved so its "unusable." You might even see these comments on HN, a place full of tech experts.

What I'm getting at here is that the efficient market hypothesis is clearly false and market participants are clearly not rational (or at least based on the conventional -- economic -- definitions)

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#137

Earlier quoted context omitted.

The only truly successful commercial use of SDXL I know of is by NovelAI. Said company appears to have used an 256xH100 cluster to finetune it to produce anime art. Open source efforts to produce a similar model seem to have failed due to the extreme compute requirements for finetuning. For example, Waifu Diffusion using 8XA40[0] have not managed to bend SDXL to their will after potentially months of training. If you…

What do you mean by extreme requirements? There's lots of SDXL fine tunings available at civit, like https://civitai.com/models/119012/bluepencil-xl for anime. The relevant discords for models/apps are full of people doing this at home. Or are you looking at some very specific definition / threshold for fine tuning here?

There is something I find rather hard to communicate about the difference between these models on civitai and what I think a competent model should be able to do.

I'd describe them like "a bike with no handlebars" because they are incredibly difficult to steer to where you want.

For example if you look at the preview images like this one: https://civitai.com/images/3615715

The model seems to have completely ignored a good 35% of the text input, most egregiously I find the (flat chest:2.0), the parenthesis denoting a strengthening of that specific part of the prompt. The values I see people use with good general models range from 1.05~1.15. 2.0 in comparison is an extremely large value, that ended up _still not working at all_, if you take a look at the actual image.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#138
post #112

Earlier quoted context omitted.

Seems like a great place to mention the MapQuest API. I tested it against 4 other services(inc. Google's) on 200 manually verified geocoding/reverse-geocoding tasks. Google managed a 92% whereas MapQuest scored 99%. If you are doing geocoding or reverse-geocoding MapQuest outperforms Google Maps quite magnificently(YMMV). The cost is also lower and there are plans where you can keep the data. The funny thing being th…

What were the others? Google Maps, MapBox, OSM are the three that come to my mind.

Waze? But my hyperbole aside ("only one"), how many people do you know use these other platforms? Colloquially people call things monopolies if they have sufficient market share, not absolute (which nearly never exists), or market collusion exists (e.g. ISPs, airlines, oil). People even do it for 2 dominating companies (e.g. Coke + Pepsi only controls 71% of market share (46.3+24.7)). Because the truth is that monopolies aren't always bad, they are just dangerous because they have so much weight that they can perform abusive tactics and that's the thing we actually care about. My whole point is that this gets to be a very sticky situation when the product is the market share (the more people that use Google Maps the better Google Maps gets. But maybe social networks are a clearer example).

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#139

Earlier quoted context omitted.

> Porn. It's always porn. I've been surprised at the explosion of porn. Well, not actually. Automatic1111 made that easy and anyone that CivitAI knows all too well what those models are being used for. I mean when you give teenagers the ability to undress their crushes[0] what do you think is going to happen (do laws adequately protect people (kids)? Can they? Will this force a shift towards actually chasing producer…

>do laws adequately protect people (kids)? Can they? Will this force a shift towards actually chasing producers, distributors, and diddlers? It's extremely complicated. Actual CSAM is very illegal, and for good reason. However, artistic depictions of such are... protected 1st Amendment expression[0]. So there's an argument - and I really hate that I'm even saying this - that AI generated CSAM is not prosecutable, as…

> I'm disgusted, but not surprised, that teenage kids are generating CSAM like this. Even before we had diffusion models, we had GANs and deepfakes, which were almost immediately used for generating shittons of nonconsensual porn

I think the big difference now is that 1) it's much easier to do now, and 2) the computational requirements and (more importantly) technical skills have dramatically dropped.

We should also be explicitly aware that deep fakes are still new. GANs in 2014 were not creating high definition images. They were doing fuzzy black and white 28x28 faces, poorly, and 32x32 color images that if you squint hard enough you could see a dog (https://arxiv.org/abs/1406.2661). MNIST was a hard problem at that time and that's 10 years. It took another 4 years to get realistic faces and objects (https://arxiv.org/abs/1710.10196) (mind you, those images are not random samples), another year to get to high resolution, and another 2 to get to diffusion and another 2 before those exploded. Deep fakes were really only a thing within the last 5 years and certainly not on consumer hardware. I don't think the legal system moves much in 10 years let alone 5 or 2. I think a lot of us have not accurately encoded how quickly this whole space has changed. (image synthesis is my research area btw)

I'm not surprised that these teenagers in a small town did this. But the fact that all those adjectives exist in that order is distinct. Discussions of deep fakes like that Tom Scott video were barely a warning (5 years is not a long time). It quickly went from researchers thinking it can happen in the next decade and starting discussions to real world examples making the news in under their prediction time (I don't think anyone expected how much money and man hours would be dumped into AI).

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#140

Earlier quoted context omitted.

>do laws adequately protect people (kids)? Can they? Will this force a shift towards actually chasing producers, distributors, and diddlers? It's extremely complicated. Actual CSAM is very illegal, and for good reason. However, artistic depictions of such are... protected 1st Amendment expression[0]. So there's an argument - and I really hate that I'm even saying this - that AI generated CSAM is not prosecutable, as…

> AI generated CSAM is not prosecutable This is true, though "AI CSAM" is an oxymoron. There is no abuse in the creation of such works, and such it is not abuse material, unless of course real children are involved.

I get your argument, but there are definitely laws about cartoon underage characters. Agree or disagree the difference is that today you don't need to be a highly skilled artist to make something that people are going to fap to. (I definitely agree priority should be focused on physical abuse and the people making the content, but this whole subject is touchy).
Post reply on HN