1) Write text to describe the problem 2) Generate an image Y that encodes that information 3) Parse that image Y to do X
Example: Y = blueprint, X = Constructing a building with that blueprint
231–240 of 481 posts
1) Write text to describe the problem 2) Generate an image Y that encodes that information 3) Parse that image Y to do X
Example: Y = blueprint, X = Constructing a building with that blueprint
Earlier quoted context omitted.
Wow, I am blown away. Some of these clips are really good! I love the Arabic Gospel one. John and George would have loved this so much. And the fact that you can make things that sound good by going through visual space feels to me like the discovery of a Deep Truth, one that goes beyond even the Fourier transform because it somehow connects the aesthetics of the two domains.
I can simultaneously burst a bubble and provide fuel for more -- the alignment of the intrinsic manifolds of different domains has been an interesting research topic for zero shot research for a few years. I remember seeing at CVPR 2018 the first zero shot...classifier, I think? That if I recall correctly trained in two domains that were automatically basically aligned with each other enough to provide very good zero…
Earlier quoted context omitted.
> composition's lacking, too. It's loop-y. Well no wonder, it has absolutely no concept of composition beyond a single 5s loop, if I understand correctly. > It absolutely sucks at cymbals, though. Everything sounds like realaudio :) > It could probably do bad modern production fairly well even now :) exaggeration, but not much, when stuff is really overproduced it starts to get way more indistinct, and this can do in…
No. And I suspect this will always have phase smearing, because it's not doing any kind of source separation or individual synthesis. It's effectively a form of frequency domain data compression, so it's always going to be lossy. It's more like a sophisticated timbral morph, done on a complete short loop instead of an individual line. It would sound better with a much higher data density. CD quality would be 220500 s…
The timbral qualities of the posted samples remind me of some of the stuff I heard from Aphex Twin, like Alberto Balsalm. Not accessible by a long shot but definitely human
This is the most exciting thing I’ve seen in ages as it shows we may be on the verge of the next wave of new technology in music that will allow all sorts of weird and wonderful new styles to emerge. I can’t wait to see what these tools can do in the hands of artists as they become more mainstream.
The vocals in these tracks are so interesting. They sound like vocals, with the right tone, phonemes. and structure for the different styles and languages but no meaning. Reminds me of the soundtrack to Nier Automata which did a similar thing: https://youtu.be/8jpJM6nc6fE
They even talk specifically about about applying stable diffusion and spectrograms.
Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…
- No new result available, looping previous clip
- Uh oh! Servers are behind, scaling up
I hope Vercel people can give you some free credits to scale it up.
This is huge. This show me that Stable Diffusion can create anything with the following conditions: 1. Can be represented as as static item on two dimensions (their weaving together notwithstanding, it is still piece-by-piece statically built) 2. Acceptable with a certain amount of lossiness on the encoding/decoding 3. Can be presented through a medium that at some point in creation is digitally encoded somewhere. Th…