Live data from Hacker News

PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

stablediffusionweb.com

11–20 of 28 posts

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#11

This appears to be work sponsored by Huawei.

I suppose it won't work for generating images of Winnie the Pooh then. But seriously, it's open-source, so it hardly matters.

It does work, very well actually. On the other hand, Dalle-3 refused, with:

> I can create an image for you, but I need to modify your request to avoid depicting specific public figures or copyrighted characters.

It took effort for even "Chinese leader".

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#12
This has problems usually not seen with current systems. It's produced human characters with one thick leg and one thin leg. Three legs of different sizes. Three arms.

It can do humans in passive poses, but ask for an action shot and it botches it badly. It needs more training data on how bodies move. Maybe load it up with stills from dance, martial arts, and sports.

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#14
post #5
post #4

The source code license is AGPL-3.0 license. Perfect for these kinds of models: https://github.com/PixArt-alpha/PixArt-alpha

hm. so actually it's not compatible with commercial projects then.

It's not compatible with the idea that you can just use the project for commercial purposes without giving some form of contribution back. That's a good outcome imho.

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#15
post #10
post #5

Earlier quoted context omitted.

hm. so actually it's not compatible with commercial projects then.

> hm. so actually it's not compatible with commercial projects then. Perfectly compatible: just keep all modifications to the original code in public. AGPL does not mean that all code from your company must be open-source, just whatever is in the same binary/program as the AGPL one.

Maybe I need to read the AGPL again. I remember interpreting it as saying that if your service was dependant on the component then the license applies to the code for the rest of the service.

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#16
post #12

This has problems usually not seen with current systems. It's produced human characters with one thick leg and one thin leg. Three legs of different sizes. Three arms. It can do humans in passive poses, but ask for an action shot and it botches it badly. It needs more training data on how bodies move. Maybe load it up with stills from dance, martial arts, and sports.

How is that not seen with current systems? Have you never used a Stable Diffusion base model?

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#17
post #16
post #12

This has problems usually not seen with current systems. It's produced human characters with one thick leg and one thin leg. Three legs of different sizes. Three arms. It can do humans in passive poses, but ask for an action shot and it botches it badly. It needs more training data on how bodies move. Maybe load it up with stills from dance, martial arts, and sports.

How is that not seen with current systems? Have you never used a Stable Diffusion base model?

I’d say SD works quite well at the macro scale, comparatively; while the artifacts the GP mentioned still occur with SD and to a lesser extent Midjourney and DALL-E, they’re much less of a common occurrence in comparison.

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#18
post #11

Earlier quoted context omitted.

I suppose it won't work for generating images of Winnie the Pooh then. But seriously, it's open-source, so it hardly matters.

It does work, very well actually. On the other hand, Dalle-3 refused, with: > I can create an image for you, but I need to modify your request to avoid depicting specific public figures or copyrighted characters. It took effort for even "Chinese leader".

Interesting, considering Winnie the Pooh is in the public domain now. I wonder how they determine the copyright status of a character.

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#19
post #18
post #11

Earlier quoted context omitted.

It does work, very well actually. On the other hand, Dalle-3 refused, with: > I can create an image for you, but I need to modify your request to avoid depicting specific public figures or copyrighted characters. It took effort for even "Chinese leader".

Interesting, considering Winnie the Pooh is in the public domain now. I wonder how they determine the copyright status of a character.

Fairly confident they just ask GPT-4 to do its best not to serve any requests that might violate copyright in the system prompt.

Re: PixArt-α:A New Open-Source Text-to-Image Model Challenging SDXL and Dalle·3

#20
post #15
post #10

Earlier quoted context omitted.

> hm. so actually it's not compatible with commercial projects then. Perfectly compatible: just keep all modifications to the original code in public. AGPL does not mean that all code from your company must be open-source, just whatever is in the same binary/program as the AGPL one.

Maybe I need to read the AGPL again. I remember interpreting it as saying that if your service was dependant on the component then the license applies to the code for the rest of the service.

Please, do read it. Please, post it here if you find something that backs your recollection.
Post reply on HN