Live data from Hacker News

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

blog.comfy.org

31–40 of 100 posts

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#31
post #26

On one hand: impressive. On the other hands aesthetically it all looks painfully bland and generic.

I agree. But is that the model or the prompt? I remember reading a report where people running AI-model instagram account were using insanely long and detailed prompts about the setting, lighting, makeup, pose, disposition, clothing, etc. about their models. Presumably with some reference image of the face / body to remain consistent across images. It‘s not clear to me whether a sufficiently detailed prompt can gener…

That would be the prompt. With the right assistance from Qwen3.5/Ornith I was able to achieve some amazing results. Unfortunately due to their licensing I'm not allowed to use it in the USA, so I had to halt testing.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#32
post #7

The mouse render is surprisingly good. Several of those clips stood out to be a pretty big leap in terms of current SOTA models. The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering. I feel like for a good while now we'll transition into a process that uses traditional "close-u…

This is still about a year and a half behind Seedance 2.0/ Seedance 2.5

But it represents a coming price pressure that will face the leading foundation models. Open source will prevent runaway costs.

Moreover, it prevents the hair-trigger platform safety checkers from shutting down creative work. Video models are notoriously bad at shutting down a huge number of requests.

Creatives will prefer to work on cloud or private GPU clusters. Waiting 10 minutes for a few seconds of 480p is unacceptable. Hobbyists will have fun, but most actual production work is happening in the cloud.

Artist's time is worth money, and they like to spin up dozens of concurrent generations at a time to more quickly explore the generation state space and make progress on completing work.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#33
post #10

Earlier quoted context omitted.

"Regions such as the EU, UK, South Korea, and the US are currently developing or enforcing AI-related regulations that may have specific implications for generative video models" You just have to pinkie promise you won't make disney mad and they will send you a licence https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/Q...

I do work animations for fun and internal use only (mostly jokes). I MAY reach out to them.

Worth if you plan to use it for production use. If you use it for personal or as mockup, then what they don't know won't hurt them? ;)

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#34
post #18
post #8

Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.

I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics. Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?

I've had great results on electronics with Claude and Gemini.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#36
post #18
post #8

Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.

I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics. Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?

Suck at what aspect of (analog) electronics specifically? Not contradicting the claim, just want to understand it.

I have not tested yet, but I suspect that LLMs with a harness that can execute code can do SPICE simulations rather ok these days? I have seen MCPs for measurement equipment also, maybe they can even close the physical loop?

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#37
post #8

Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.

Huh, interesting. I tried to generate a 10 second 1080p clip on a bigger machine and the results were quite poor. Unusable for anything, in fact.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#38

> We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality. Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as w…

Whoah, could this mean we can treat layers like a jpg, where we come up with a formula that estimates the weight values of a layer instead of storing all of the weights?

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#39
post #26

On one hand: impressive. On the other hands aesthetically it all looks painfully bland and generic.

I agree. But is that the model or the prompt? I remember reading a report where people running AI-model instagram account were using insanely long and detailed prompts about the setting, lighting, makeup, pose, disposition, clothing, etc. about their models. Presumably with some reference image of the face / body to remain consistent across images. It‘s not clear to me whether a sufficiently detailed prompt can gener…

I mean, even the demo prompts on the ComfyUI page aren't adhered to by the model. From the first prompt, one of the four lines:

> TRANSITION: a violent WHIP PAN off the rooftop that SMEARS the floating words away with it, motion-streaked —

And the video just didn't do any of that transition at all, it just replaced it with a cut. If you look at the rest of the prompts, you'll find similar lines that are just totally ignored. Except maybe the mouse one, I didn't see anything wrong with that off the bat.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#40

I've said it before and I'll say it again, human directors are still valuable, as they use AI video editing tools to generate the shots they want and put them together in a cohesive way. Previously they might've used film and actors but if they can just prompt the AI (or create workflows as seen with ComfyUI) then they arrange them together just like how an EDM producer doesn't actually play the instruments but inste…

Seedance 2.5 just came out and it is incredible,significantly better than this for a lot of cases. This one is the latest _free_ video generator.

But as far as your composition tools, they are already available.

Post reply on HN