The main scalability issue with this is that the products you've found will quickly go out of stock, so as you build up a collection of outfits you'll face an increasing burden going back and checking all of the products to see if they're in stock (or risk damaging your brand).
I worked on something similar back in 2008. We were looking at ways of monetising our visual similarity engine. We could mark a set of query products for each outfit and return a selection of products that were both similar and in-stock and give the customer the option of filtering by price range or whatever.
There were some nice challenges in there, like processing gigabytes of retailer feeds as rapidly as possible looking for new items, standardising various huge feeds without using up developer time, product deduplication, image feature extraction, designing the indexing method (we ended up using the Visual Words technique with a custom distributed Lucene inverted index as Solr didn't support partitions at the time). It was a really fun project... and I've drifted far enough off topic that I'm going to finish up.
The tech was pretty solid (and replicable if you can get someone decent to do the CBIR piece) but we ran out of runway.