Earlier quoted context omitted.
It's statistical prediction. LLMs do not "understand" the world by definition. Ask an image generator to make "an image of a woman sitting on a bus and reading a book". Images will be either a horror show or at best full of weird details that do not match the real world - because it's not how any of this works. It's a glorified auto-complete that only works due to the massive amounts of data it is trained on. Throw i…
Why do people say stuff like this that is so demonstrably untrue? SD and GPT4 do not exhibit the behavior described above and they're not even new.
Here's 1.5 EMA https://imgur.com/mJPKuIb
Here's 2.0 EMA https://imgur.com/KrPVUGy
No negatives, no nothing just the prompt. 20 steps of DPM++ 2M Karras, CFG of 7, seed is 1.
Can we make it better? Yeah sure, here's some examples: https://imgur.com/Dmx78xV, https://imgur.com/HBTitWm
But I changed the prompt and switched to DPM++ 3M SDE Karras
Positive: beautiful woman sitting on a bus reading a book,(detailed [face|eyes],detailed [hands|fingers]:1.2),Tokyo city,sitting next to a window with the city outside,detailed book,(8k HDR RAW Fuji film:0.9),perfect reflections,best quality,(masterpiece:1.2),beautiful
Negative: ugly,low quality,worst quality,medium quality,deformed,bad hands,ugly face,deformed book,bad text,extra fingers
We can do even better if we use LoRAs and textual inversions, or better checkpoints. But there's a lot of work that goes into making really high quality photos with these models.
Edit: here is switching to Cyberrealistic checkpoint: https://imgur.com/gFMkg0J,
And here's adding some LoRAs, TIs, and prompt engineering:
https://imgur.com/VklfVVC (https://imgur.com/ZrAtluS, https://imgur.com/cYQajMN), https://imgur.com/ci2JTJl (https://imgur.com/9tEhzHF, https://imgur.com/4Ck03P7).
I can get better, but I don't feel too much like it just to prove a point.