It feels cheaper by orders of magnitude right now per-second. I'm not sure how long that will continue, but until then, that's pretty important to me, especially when video generation is so imprecise, and thus more of a "rough draft" creative tool than a "final shot" tool.
Quite often my workflow looks like: "Grab this frame from which I'd like a continuation" -> "Upload back to Grok on web" -> "Generate until happy" -> "Download and use".
So a feature request (or I'll just fork next time I'm doing something with AI video) is simply a browser tab alongside the generation tab, literally just to make that window swapping easier, and ideally with 'download' pointing to a sensible location to avoid file replication.
This feels closest to how I use AI video, which is just kind of like an infinitely extensible library tab on my video editing software. I know you said 'no webview' - I don't know Swift well enough to know if that excludes this feature.
There's probably a bunch of ways to improve my specific flow once there's a browser within the UI, especially with MCP integration. Much to consider.
And overall it looks fun, and I'm excited to give it a try. It's great this has beat detection specifically, because that's non-negotiable in a video editor for me when doing AI stuff, as the medium remains more suited to non-dialogue work.
Similarly, the CapCut 'AI video editor' that's most similar to the 'MCP' feature is unusably chaotic, so I'm interested to see how well models do in your architecture. I can understand why you've found structures like 'beat' or 'transcription' massively aid it.
Final Cut export - if it isn't supported by what you've already done (I don't move stuff between editors much!) - is really useful to me personally also.
Best of luck! Excited to test it! Thanks for sharing.