How does this compare with sora (pro)?
Sora, the image model (gpt-image-1), is phenomenal and is the best-in-class.
I can't wait to see where the new Imagen and Veo stack up.
171–180 of 570 posts
How does this compare with sora (pro)?
Sora, the image model (gpt-image-1), is phenomenal and is the best-in-class.
I can't wait to see where the new Imagen and Veo stack up.
It finally feels like the professional tools have greatly outpaced the open source versions. While wan and hunyuan are solid free options, the latest from Google and Runway have started to feel like a league above. Interestingly it feels like the biggest differentiator is editing tools - ability to prompt motion, direction, cuts, or weaving in audio, rather than just pure ability to one shot. These larger companies a…
I think open source still has an important advantage in the pro environment despite being less convenient, and it's the possibility of adding things in between the generation process like control net, and custom loras with new concepts or characters. Plus in local generation you're not limited by the platform moderation that can be too strict and arbitrary and fail with the false positives. Yes comfy UI can be intimi…
Earlier quoted context omitted.
The amount of gatekeeping I see when this topic is brought is outstanding! Why can't people be happy that more individuals would be soon able to create freely in a more accessible way? Personally I can't wait to see the new creative doors ai will open for us!
How is the requirement to use a computer and maybe pay a cloud subscription in the long term more accessible than other kinds of art? Which individuals are gatekept exactly? Before you bring up disabled people (as often happens when the term accessibility is used), know that many of them are not happy to be used as a shield for this without ever being asked and would rather speak for themselves. I've tried AI image g…
Other disapproval comes from different emotional places: a retreading of ludditism borne out of job insecurity, criticism of a dystopia where we've automated away the creative experience of being human but kept the grim work, or perceptions of theft or plagiarism.
Whether AI has worked well for you isn't just irrelevant, but contrarian in the face of clear and present value to a lot of people. You can be disgusted with it but you can't claim it isn't there.
Earlier quoted context omitted.
It's being redefined in such a way that 2-3 very large entities get to hold the means of production. It's a very convenient redefinition for them.
[flagged]
This view also aligns with how generative AI is marketed – it's a way to accelerate realization, not a way to focus on the act of crafting.
That said, outcome-first thinking does run the risk of disconnection, and our current culture is all about disconnection.
Earlier quoted context omitted.
So deliberately writing a prompt that meticulously describes how a generated photo would look like isn't creative, but pushing a button for a machine to take the photo for you is??!! If anything, it's the way around! Of course that's not what I believe, but let's not limit the definition of what creativity based on historical limitations. Let's see what the new generation of artists and creators will use this new cap…
Your meticulous prompt is using the work of thousands of experts, and generating a mashup of what they did/their work/their commitment/their livelihood. Their placement of books. Their aesthetic. The collection of cool things to put into a scene to make it interesting. The lighting. Not yours. Not from you/not from the AI. None of it is yours/you/new/from the AI. It's ALL based underneath on someone else's work, some…
After doing some testing, Imagen 4 doesn't score any higher than Imagen 3 on my comparison chart, approximately ~60% prompt adherence accuracy. https://genai-showdown.specr.net
The winning image entry for "The Yarrctic Circle" by OpenAI 4o doesn't actually wields a cutlass. It's very aesthetically pleasing, even though it's so wrong in all fundamental aspects (perspective is nonsensical and anatomy is messed up, with one leg 150% longer than the other, ...). It's a very interesting resource to map some of the limits of existing models.
Earlier quoted context omitted.
It's being redefined in such a way that 2-3 very large entities get to hold the means of production. It's a very convenient redefinition for them.
[flagged]
Step two is... sure, every pleb can now create art.
That devalues art. More than that, that makes for a "winner takes all" marketplace. So even fewer people than now benefit from it. More than that, guess who wins out: middlemen, the marketplace owners.
Read the Black Swan by Nassim Taleb, especially the chapters about Extremistan and Mediocristan. Basically every time we invent something that scales and unlocks something for a great amount of people, we commoditize it and the quality of life for the average person in that field goes down while the leeches, pardon me, the middle men, are the only ones that become constantly rich, after the initial struggle to achieve market dominance (so when the market matures).
>>models create, empowering artists to bring their creative vision Interesting logic the new era brings: something else creates, and you only "bring your vision to life", but what it means is left for readers questioning, your "vision" here is your text prompt? Were at a crossroads where the tools are powerful enough to make the process optional. That raises uncomfortable questions: if you don’t have to create anymor…
It finally feels like the professional tools have greatly outpaced the open source versions. While wan and hunyuan are solid free options, the latest from Google and Runway have started to feel like a league above. Interestingly it feels like the biggest differentiator is editing tools - ability to prompt motion, direction, cuts, or weaving in audio, rather than just pure ability to one shot. These larger companies a…
The Tencent Hunyuan team is cooking.
Hunyuan Image 2.0 [1] was announced on Friday and it's pretty amazing. It's extremely high quality text-to-image and image-to-image with millisecond latency [2]. It's so fast that they've built a real time 2D drawing canvas application with it that pretty much duplicates Krea's entire product offering.
Unfortunately it looks like the team is keeping it closed source unlike their previous releases.
Hunyuan 3D 2.0 was good, but they haven't released the stunning and remarkable Hunyuan 3D 2.5 [3].
Hunyuan Video hasn't seen any improvements over Wan, but Wan also recently had VACE [4], which is a multimodal control layer and editing layer. The Comfy folks are having a field day with VACE and Wan.
[1] https://wtai.cc/item/hunyuan-image-2-0
[2] https://www.youtube.com/watch?v=1jIfZKMOKME&t=1351s
[3] https://www.reddit.com/r/StableDiffusion/comments/1k8kj66/hu...
Earlier quoted context omitted.
[flagged]
If your focus is to solve the problem, then it makes sense to treat the process as secondary. The tools are just means to an end. This view also aligns with how generative AI is marketed – it's a way to accelerate realization, not a way to focus on the act of crafting. That said, outcome-first thinking does run the risk of disconnection, and our current culture is all about disconnection.
Build out your app idea first with replit -> Then export the codebase into your computer -> Run claude code on it and ask it to scan all the files and describe the tech stack to you and how it operates while giving you all the major components you need to learn to understand it with youtube channel and book recommendations for each topic + work exercises -> Use perplexity deep research once a week to further research every topic as you start to learn them
If you’re a busy man/woman make gumloop or lindyai workflow to check your calendar and pack in timeslots to do all of this learning, and then auto send you worksheets via email as homework to test you skills
All of this for a price of 1/15th of a college degree (not even an expensive college)
This is not hypothetical conjecture I do this daily.
So everyone has now 1) Low cost access to build stuff with one prompt to realise the value of tools 2) A personal tutor that can then help you scour the depths of the craft and force you to practice and learn deeply now with your added motivation of knowing what’s possible with building stuff
So it has the potential to connect us more too, it’s upto humans to choose whether they do at the end tho. That is their liberty.