Earlier quoted context omitted.
Do you have link for the VS Community terms you're describing? What I've found is directly contradictory: "Any individual developer can use Visual Studio Community to create their own free or paid apps." From https://visualstudio.microsoft.com/vs/community/
Enterprise organizations are not allowed to use VS Community for commercial purposes: > In enterprise organizations (meaning those with >250 PCs or >$1 Million US Dollars in annual revenue), no use is permitted beyond the open source, academic research, and classroom learning environment scenarios described above.
Stable Video Diffusion
301–310 of 316 posts
Re: Stable Video Diffusion
#302In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
Adobe is doing some great work here in my opinion in terms of building AI tools that make sense for artist workflows. This "sneak peak" demo from the recent Adobe Max conference is pretty much exactly what you described, actually better because you can just click on an object in the image and drag it. See video: https://www.adobe.com/max/2023/sessions/project-stardust-gs6...
Re: Stable Video Diffusion
#303Earlier quoted context omitted.
> my wife walks up to my computer and makes a copy of the data As you agreed to in our contract, you now need to compensate me for the damage caused by your failure to prevent unauthorized third-party access. Of course you're free to attempt to recover the sum you have to pay me from your wife. > The output of a program is rarely copyrightable by the author (as opposed to the user). The author of the program can make…
Okay. Now put yourself in the position of Microsoft, using this scheme for Windows. We'll pretend real copyright doesn't exist, and we've got your hair-brained scheme. This is how it plays out: 1) You have a $1T product. 2) My wife leaks it, or a burglar does. I am a typical consumer, with say, a $20k net worth. You have two choices: 1) Sue me, recover $20k, and be down $1T (minus $20k, plus litigation fees), and get…
It seems to me that the cases in the article you linked involved the author of the program arguing that their copyright automatically extended to the output without any extra contractual provisions concerning copyright assignment, so I don't think they can be used as precedent regarding the enforceability of such clauses.
Re: Stable Video Diffusion
#304Earlier quoted context omitted.
Oh, I don't know about that. Working in feature film animation, studios have gargantuan model libraries from current and past projects, with a good number (over half) never used by a production but created as part of some production's world building. Plus, generative modeling has been very popular for quite a few years. I don't think getting more 3D models then they could use is a real issue for anyone serious.
Where can you find those? I'm in the same situation as him, I've never heard of a 3d dataset better than objaverse XL. Got a public dataset?
I've not worked in VFX for a while, but when I did the modeling departments at multiple studios had giant libraries of completed geometries for every project they ever did, plus even larger libraries of all the pieces and parts they use as generic lego geometry whenever they need something new.
Every 3D modeler I know has their own personal libraries of things they'd made as well as their own "lego sets" of pieces and parts and generative geometry tools they use when making new things.
Now this is just a guess, but do you know anyone going through one of those video game schools? I wager the schools have big model libraries for the students as well. Hell, I bet Ringling and Sheridan (the two Harvards of Animation) have colossally sized model libraries for use by their students. Contact them.
Re: Stable Video Diffusion
#305In the video towards the bottom of the page, there are two birds (blue jays), but in the background there are two identical buildings (which look a lot like the CN Tower). CN Tower is the main landmark of Toronto, whose baseball team happens to be the Blue Jays. It's located near the main sportsball stadium downtown. I vaguely understand how text-to-image works, and so it makes sense that the vector space for "blue j…
> sportsball This is not the flex you think it is. You don't have to like sports, but snarking on people who do doesn't make you intellectual, it just makes you come across as a douchebag, no different than a sports fan making fun of "D&D nerds" or something.
Re: Stable Video Diffusion
#306Earlier quoted context omitted.
> Back in the mid 90s to 2010 or so, graphical improvements were hailed as photorealistic Whenever I saw anybody calling those graphics "photorealistic", I always had to roll my eyes and question if those people were legally blind. Like, c'mon. Yeah, they could be large leaps ahead of the previous generation, but photorealistic? Get real. Even today, I'm not sure there's a single game that I would say has photo-reali…
> Even today, I'm not sure there's a single game that I would say has photo-realistic graphics. Looking just at the videos (because I don't have time to play the latest games any more and even if I did it's unreleased), I think that "Unrecord" is also something I can't distinguish from a filmed cinematic experience[0]: https://store.steampowered.com/app/2381520/Unrecord/ Though there are still caveats even there, as…
IMO, though, the lighting in the indoor scenes is just not quite right. There's something uncanny valley about it to me. When the flashlight shines, it's clearly still a computer render to my eyes.
The outdoor shots, though, definitely look flawless.
Re: Stable Video Diffusion
#307Earlier quoted context omitted.
No I think we’re actually close. My source is I’m working on this problem and the incredible progress of our tiny 3 person team at drip.art ( http://api.drip.art ) - we can generate a lot of frames that are consistent, and with interpolation between them, smoothly restyle even long videos. Cross-frame attention works for most cases, it just needs to be scaled up. And that’s just for diffusion focused approaches like…
The demo clips on the site are cool, but when you call it a "solved problem," I'd expect to see panning, rotating, and zooming within a cohesive scene with multiple subjects.
Re: Stable Video Diffusion
#308Re: Stable Video Diffusion
#309Earlier quoted context omitted.
It doesn't have to be enforceable. This licensing model works exactly the same as Microsoft Windows licensing or WinRAR licensing. Lots and lots of people have pirated Windows or just buy some cheap keys off Ebay, but no one of them in their sane mind would use anything like that at their company. The same way you can easily violate any "non-commercial" clauses of models like this one as private person or as some tin…
I've heard companies also intentionally do not go after individuals pirating software e.g., Adobe Photoshop - it benefits them to have students pirate and skill up on their software and then enter companies that buy Photoshop because their employees know it, over locking down and having those students, and then the businesses, switch to open source.
One smart thing they did was they'd check the online job listings and if a firm advertised for needing AutoCAD experience they'd check their licenses. I knew firms who got calls from Autodesk legal the DAY AFTER posting an opening.
Re: Stable Video Diffusion
#310Earlier quoted context omitted.
But surely you wouldn't try to emit that format directly, but rather some higher level scene description? Or even just a set of instructions for how to manipulate the UI to create the imagined scene?
It sure feels weird to me as well, that GenAI is always supposed to be end-to-end with everything done inside NN blackbox. No one seems to be doing image output as SVG or .ai.
LLMs are good at writing copy that sounds accurate and creative enough, and there are known techniques to improve that (such as generating an outline first, then generating each section separately). If you then give them a list of templates, and written examples of what they are used for, the LLM is able to pick one that's a suitable match. But this is all just probability, there's no real creativity here.
Earlier this year I played around with trying to have GPT-3 directly output an SVG given a prompt for a simple design task (a poster for a school sports day), and the results were pretty bad. It was able to generate a syntantically coreect SVG, but the design was terrible. Think using #F00 and #0F0 as colours, placing elements outside the screen boundaries, layering elements so they are overlapping.
This was before GPT-4, so it would be interesting to repeat that now. Given the success people are having with GPT-4V, I feel that it could just be a matter of needing to train a model to do this specific task.