SD outputs have an "uncanny valley" type of quality to them. You just KNOW when an image is from SD. And I have looked at getting started with SD, but the requirements and setup and +-prompting "language" just kind of turned me off the whole thing. Whereas with DALL E you can get some hyper-realistic images from it with very little effort using plain human language. I guess my point is to ask whether SD is worth both…
No, you know when a beginner generated an image in Stable Diffusion. With enough skill and attention, you will not.
Sure, there is a learning curve and it takes more time to get to a good result. But in turn, it gives you control far beyond what the competition can offer.