Live data from Hacker News

SDXL Turbo: A Real-Time Text-to-Image Generation Model

stability.ai

111–120 of 157 posts

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#111
post #106
post #98

Earlier quoted context omitted.

If these creations inspire such violent disgust, then it's likely that you perceive them as authentic art. If AI images were devoid of meaning or value, they wouldn't have sparked such passion. You cannot claim ownership over culture, nor can AI. Culture is a collaborative process, and no one can barricade themselves from the input of others. Artists using AI are simply exercising their right to contribute to the col…

This is, imo, an extremely naive take. You claim culture is a collaborative process, yet AI only takes from the communities that produce art. It gives nothing back. You claim AI produces culture, but all it does is atomize our society, promising personal yet meaningless experiences for everyone. There's no shared culture if everyone is just consuming individualized streams of content. It's simultaneously homogenizing…

I don't try to swim against the current. But you're welcome to do it.

What do you mean AI doesn't give back? It serves everyone and gives back everything it can create. Artists are the number one users here, and they will unlock the AI skills better than regular people playing around.

What individualized streams? you mean like imagination, where everyone of us has their own "individualized stream"? AI art is augmented imagination. No obstacle in sharing, in fact it's easier now. You don't need to be an artist to create depictions of your imaginations, and sharing a generated JPG is much easier than drawing it by hand.

> yet you continue to use the labor of others without permission

That's how culture works. The artist who never took inspiration from the cultural environment should throw the first stone. Pablo Picasso is widely quoted as having said that “good artists borrow, great artists steal.”

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#112

Earlier quoted context omitted.

The EU views Google as a tobacco company. The last thing they want to do is bankrupt them with fees. They want to milk Google - big tech generally - for tax revenue. And besides, it'd take $100 billion per year in fees, which is never going to happen. Meanwhile Google keeps getting bigger year after year (they have nearly doubled in size in four years, up to $300b in sales now) and Bing has made zero headway despite…

> The US breaking Google up also won't end the monopoly money. It's also weird because modern economics with tech has created a space that creates a lot of natural monopolies. Momentum is very powerful and it's the reason silicon valley companies will run at a loss for years creating a userbase. Trick is to keep them (or sell before buyer starts charging). There's what, 2 map companies and only one of them is in high…

Seems like a great place to mention the MapQuest API. I tested it against 4 other services(inc. Google's) on 200 manually verified geocoding/reverse-geocoding tasks. Google managed a 92% whereas MapQuest scored 99%.

If you are doing geocoding or reverse-geocoding MapQuest outperforms Google Maps quite magnificently(YMMV). The cost is also lower and there are plans where you can keep the data.

The funny thing being that the C levels didn't give a shit about the results and went ahead with Google Maps anyway. So your point remains correct.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#113
post #90
post #76

Earlier quoted context omitted.

The so called "AI People" built the entire architecture, something people didn't think was possible at the scale and quality a year ago, and the matter of "artists should get whatever they want" because it trained on their works isn't the point. Diffusion Models don't rip parts of pictures together, they happen to be trained to make art out of noise, finding patterns in art. same things happening with LLM's in court…

These models can't exist without the training sets. Their value is entirely derived from existing data. The ml architecture does not matter at all. Sure, throw enough compute and data at a problem, do a little parallelization, and you can extract plenty of patterns. Does that mean the ml engineers understand art? Or are they just using glorified brute force to alienate people who actually make things from their labor…

Do humans understand art? Does it matter? How do humans learn to create art? By looking at other peoples art, for the most part. Those other artists are not compensated for this either, nor do they need to be credited. If you exactly copy another persons art style, that may be frowned upon by some, but otherwise it is of no consequence, unless you claim the work is actually made by that artist. If you believe that humans should be afforded a privilege that machines (or rather their operators) should not be afforded, make your case on why.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#114
post #107

Earlier quoted context omitted.

I can't reply to dead comment, but I'll reply to yours RE above: > You can rip off other people's work faster than ever! I have views on IP that mean I reject the ripping off premise on it's face, BUT IF I DID ACCEPT IT Who am I ripping off? What artist is being denied work if i reskin a VR world on the fly with AI? nobody was going to be painting a scene's worth of textures in a few seconds, the AI is enabling new c…

It's so interesting to me that hackernews will flag any comment that dissents to the use of this technology. Supposedly it's irrelevant to the conversation. The value of these models is derived from the training set, not the ml model. Take away the training data and the model does nothing. So who cares if you reskin your vr world on the fly with AI? I'd argue many artists whose work was ingested into these models wit…

I have the opposite view of you on the IP front (former 'pirate'), but your critique of HN flagging is spot on.

How do you interpret modern web design where everything looks like everything else? Is the next boring SPA also problematic IP usage?

These are genuine questions and not meant as an irritant. My view on IP is fairly extreme and I appreciate views that discourage such. Nothing you said about artists needing to eat moved the needle, though I can understand why it would for another person.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#115

Noncommercial use - aside from being one of my licensing pet peeves - seems to indicate that the money is drying up. My guess is that the investors over at Stability are tired of subsidizing the part of the generative AI market that OpenAI refuses to touch[0]. The thing is, I'm not entirely sure there's a paying portion of the market? Yes, I've heard of people paying for ChatGPT because it answers programming questio…

There's quite a few hosted SDXL platforms (mage.space, leonardo.ai, novel.ai, tensor.art, invoke.ai to name a few) and most consumers do not have the GPUs needed to run those models, only enthusiasts do.

It's always baffled me that stability didn't offer a competitive UI platform to use their models with, clipdrop is just bad quality and very bare-bones, and dreamstudio is pricey and still lacks most features. So this move to a new licensing strategy doesn't surprise me, it actually is somewhat comforting, as i expecting them to just stop releasing further trained models (e.g sdxl1.1 and up), and only offer those on their services (of course, that can still happen) cause how else were they going to monetize the consumers (i know they (planned to) offer custom trained/finetuned models to big corps, but that doesn't monetize consumers).

However, as most releases by stability these days, it has this feeling of close-but-no-cigar, and the recent LCM lora's might be a little slower, but these actually offer 1024^2 resolution, work with any existing lora's and finetunes (so they are usable for iterative development, unlike this turbo model, cause well, it's a different model, can't iterate on it then expect sdxl (with lora's, to a lesser extend also without) to generate a similar image) and support cfg-scale (and therefor negative prompts / prompt weighting). I suppose there's some niche market where you need all the speed you can get, but unless there's a giant leap in (temporal) consistency, that will remain niche, i don't see the mentioned real-time 3d "skinning" neither the video img-to-img (frame-to-frame) gimmicks take off with current quality and lack of flexibility. It's good research, optimizations have lots of value, but it needs quality as well.

Their recent video model is quite bad as well, especially compared to pika and runway gen-2, but well, but as with the the dalle-3 comparison one can say those are closed source and stability's offering is open.

Then we have the 3d model, close sourced, worse than luma's genie unfortunately.

The music model is nothing like suno's chirp (which might be multiple models, bark and a music model) used together), and the less said about their llm offerings the better.

Bottom line, stability needs a killer model again, they started strong with stable diffusion 1.5, took a wrong turn with 2.0 (kind of recovered by 2.1, but the damage was done), and while SDXL is't bad in a vacuum, neither was it the leap ahead that put it in front of competition like midjourney at the time, and Dalle-3 a little later, and now even a relatively small model like pixart-alpha, also opensource, can offer similar quality to what sdxl offers (with a lot of caveats, as it has been trained on so few images it just doesn't have info on many concepts). And more worrying, there's no hint of something better in the stability's pipeline. But maybe image-gen is as best as stability can get it, and they think they can make an impact pivoting in another direction or multiple directiobs, but currently, it feels a master-of-none situation.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#116

Earlier quoted context omitted.

> Porn. It's always porn. I've been surprised at the explosion of porn. Well, not actually. Automatic1111 made that easy and anyone that CivitAI knows all too well what those models are being used for. I mean when you give teenagers the ability to undress their crushes[0] what do you think is going to happen (do laws adequately protect people (kids)? Can they? Will this force a shift towards actually chasing producer…

>do laws adequately protect people (kids)? Can they? Will this force a shift towards actually chasing producers, distributors, and diddlers? It's extremely complicated. Actual CSAM is very illegal, and for good reason. However, artistic depictions of such are... protected 1st Amendment expression[0]. So there's an argument - and I really hate that I'm even saying this - that AI generated CSAM is not prosecutable, as…

Do non-consensual porn not qualify as defamation? That and obscenity laws if existed should be able to handle most hyperrealistic porn so that only speeches remain.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#117

Hugging Face released a Colab Notebook for generation from SDXL Turbo using the diffusers library: https://colab.research.google.com/drive/1yRC3Z2bWQOeM4z0FeJ0... Playing around with the generation params a bit, Colab's T4 GPU can batch-generate up to 6 images at a time at roughly the same speed as one.

thanks for sharing

I got to experience the power of current models with just 5 lines of code

the pace of change is stressing me out :)

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#118

Based on the demo, that's... incredibly fast. Literally generating images faster than I can type a prompt. They've clearly got a set seed, so they're probably caching request, but even with prompts that they couldn't possibly have cached it's within a second or so.

This isn't cached. I'm running it locally on a 4060 TI 16 GB and it's just as fast. Image gens in .6-.8 seconds. Each word or character I type is a new image gen and it's INSTANT.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#119

> A Real-Time Text-to-Image Generation Model > On an A100, SDXL Turbo generates a 512x512 image in 207ms (prompt encoding + a single denoising step + decoding, fp16), where 67ms are accounted for by a single UNet forward evaluation. Okay... so what part of this is real time? 207ms is 4.8Hz. 67ms is 14.9Hz. Isn't "real time" in graphics considered to be at least 30Hz (33ms)? And by today's standards at minimum 60Hz (1…

Considering the use-case of just interacting with a computer and typing prompts, I'd call it real-time.

Re: SDXL Turbo: A Real-Time Text-to-Image Generation Model

#120

Noncommercial use - aside from being one of my licensing pet peeves - seems to indicate that the money is drying up. My guess is that the investors over at Stability are tired of subsidizing the part of the generative AI market that OpenAI refuses to touch[0]. The thing is, I'm not entirely sure there's a paying portion of the market? Yes, I've heard of people paying for ChatGPT because it answers programming questio…

The only truly successful commercial use of SDXL I know of is by NovelAI. Said company appears to have used an 256xH100 cluster to finetune it to produce anime art. Open source efforts to produce a similar model seem to have failed due to the extreme compute requirements for finetuning. For example, Waifu Diffusion using 8XA40[0] have not managed to bend SDXL to their will after potentially months of training. If you…

They rented 8xA40 for $3.1k. That is actually kinda peanuts; I spent more on my gaming PC. I think there were kickstarter projects for AI finetunes that raised $200k before Kickstarter banned them?
Post reply on HN