Earlier quoted context omitted.
But surely you wouldn't try to emit that format directly, but rather some higher level scene description? Or even just a set of instructions for how to manipulate the UI to create the imagined scene?
I've seen this but producing Python scripts that you run in Blender, e.g. https://www.youtube.com/watch?v=x60zHw_z4NM (but I saw something marginally more impressive, not sure where though!)
Stable Video Diffusion
271–280 of 316 posts
Re: Stable Video Diffusion
#272Earlier quoted context omitted.
Also the level of fault tolerance... if your pixels are a bit blurry, chances are no one notices at a high enough resolution. If your json is a bit blurry you have problems.
You can do "constrained decoding" on a code model which keeps it grammatically correct. But we haven't gotten diffusion working well for text/code, so generating long files is a problem.
I'm not experienced enough to validate their claims, but I love the choice of languages to evaluate on:
> Python, Bash and Excel conditional formatting rules.
Re: Stable Video Diffusion
#273The rate of progress in ML this past year has been breath taking. I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the temporal consistency. Getting stable video out of stable diffusion typically involves lots of manual post processing to remove flicker.
What was the big “unlock” that allowed so much progress this past year? I ask as a noob in this area.
Re: Stable Video Diffusion
#274Earlier quoted context omitted.
> you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc. I'm always confused why I don't hear more about projects going in this direction. Controlnets are great, but there's still quite a lot of hallucination and other tiny mistakes that a skilled human would never make.
I think the bottleneck is data For single 3D object the biggest dataset is ObjaverseXL with 10M samples For full 3D scenes you could at best get ~1000 scenes with datasets like ScanNet I guess Text2Image models are trained on datasets with 5 billion samples
Re: Stable Video Diffusion
#275Re: Stable Video Diffusion
#276The rate of progress in ML this past year has been breath taking. I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the temporal consistency. Getting stable video out of stable diffusion typically involves lots of manual post processing to remove flicker.
What was the big “unlock” that allowed so much progress this past year? I ask as a noob in this area.
Re: Stable Video Diffusion
#277Earlier quoted context omitted.
> Has anyone come across a solution where model can iterate (eg, with prompts like "move the bicycle to the left side of the photo")? It feels like we're close. I feel like we're close too, but for another reason. For although I love SD and these video examples are great... It's a flawed method: they never get lighting correctly and there are many incoherent things just about everywhere. Any 3D artist or photographer…
> you'll describe something and you'll get a full 3D scene, with 3D models, source of lights set up, etc. I'm always confused why I don't hear more about projects going in this direction. Controlnets are great, but there's still quite a lot of hallucination and other tiny mistakes that a skilled human would never make.
So we could have something to convert AI-generated image output into 3D scenes without having to explicitly train the "creative" AI for that.
Probably much more viable, because the quantity of 3D models out in the wild is far far lower than that of bitmap images.
Re: Stable Video Diffusion
#278Earlier quoted context omitted.
we're working on this if you want to give it a try - dream3d.com
You should put a demo on the landing page
Re: Stable Video Diffusion
#279A seemingly off topic question, but with enough compute and optimization, could you eventually simulate “reality”? Like, at this point, what are the technical counters to the assertion that our world is a simulation?
It's purely a religious question. When humanity invented the wheel, religion described the world as a giant wheel rotating in cycles. When humanity invented books, religion described the world as a book, and God as a it's writer. When humanity invented complex mechanism, religion described the world as giant mechanism and God as a watchmaker. Then computers where invented, and you can guess what happened next.
Re: Stable Video Diffusion
#280Looks like I'm still good for my bet with some friends that before 2028 a team of 5-10 people will create a blockbuster style movie that today costs 100+ million USD on a shoestring budget and we won't be able to tell.
Geordi: "Computer, in the Holmesian style, create a mystery to confound Data with an opponent who has the ability to defeat him"