Earlier quoted context omitted.
Neither of these comments are accurate. (edit: but renegade-otter is more correct) Here's 1.5 EMA https://imgur.com/mJPKuIb Here's 2.0 EMA https://imgur.com/KrPVUGy No negatives, no nothing just the prompt. 20 steps of DPM++ 2M Karras, CFG of 7, seed is 1. Can we make it better? Yeah sure, here's some examples: https://imgur.com/Dmx78xV , https://imgur.com/HBTitWm But I changed the prompt and switched to DPM++ 3M SDE…
You kind of proved my point. Of course the "finger situation" is getting better but people handling complex objects is still where these tools trip. They can't reason about it - they just need to see enough data of people handling books. On a bus. Now do this for ALL possible objects in the world. I have generated hundreds of these - the bus cabin LOOKS like a bus cabin, but it's a plausible fake - the poles abruptly…
Yeah I did say you were more right. But it was difficult to distinguish exaggeration from actual intent. You can check my comment history of me battling the common ML mindset. I love the area of study (I'm a researcher myself) but there's a lot of problems that even in the research community a lot want to ignore. It's odd to me. It's been hilarious to watch big names claim Sora understands physics. Or people think just because it doesn't understand physics that the videos aren't still impressive and even useful.
But with how you updated your language, I think we are in a very high level of agreement. You are perfectly right: no ML model "understands" anything. GPT doesn't understand how to code and image models don't understand how to... art(?) or do physics or whatever. They don't have world models. And I'm deeply frustrated that people think a single example of a accurately acting like a world model is proof and will do gymnastics to say a single counter example isn't. A single counter does disprove a world model and to understand you need to be able to self-correct. Hallucinations are fine but "are you sure?" should be enough to get it to reconsider, not double down or just switch. We can be fooled with setups, but we laugh at ourselves quickly because we self-correct fairly easily (or rather, we can).