MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
61–70 of 100 posts
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#62Earlier quoted context omitted.
FWIW on a 6000 Pro it takes 68 seconds for 10 seconds 480p video, (cold) same demo workflow as you used. Set "megapixels" to 2.0 (1920x1088) and same video seems to take 5+ minutes, not sure if everything is right/correct at the moment. As far as I can tell, the current ComfyUI nodes don't even do compilation, and I haven't looked into what attention mechanism they're using, but I'm sure with time these durations wil…
Waiting for the “FWIW on a B200” reply
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#63Earlier quoted context omitted.
FWIW on a 6000 Pro it takes 68 seconds for 10 seconds 480p video, (cold) same demo workflow as you used. Set "megapixels" to 2.0 (1920x1088) and same video seems to take 5+ minutes, not sure if everything is right/correct at the moment. As far as I can tell, the current ComfyUI nodes don't even do compilation, and I haven't looked into what attention mechanism they're using, but I'm sure with time these durations wil…
What’s the vram usage? I’ve got a 5090 so as long as I can fit it in memory the gen times should be roughly equivalent ~10%
diffusion model: minimax_h3_fl2va_bf16.safetensors
text encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
video VAE: minimax_h3_video_vae_fp16.safetensors
audio VAE: minimax_h3_audio_vae_fp32.safetensors
Ends up at ~83GB, but they also shipped bunch of pruned + quantized versions of the diffusion model, might fit with a 5090: https://huggingface.co/Comfy-Org/MiniMax-H3Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#64Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.
FWIW on a 5080 16GB it takes 3 minutes for 10 seconds 480p video (the mouse video workflow with length changed from 5 seconds to 10 seconds)
I'm on an RTX 5090. I told it to make a 5 second 864x480 video, it's been running for over 30 minutes and is only 35% done in the SamplerCustomAdvanced step.
EDIT: Oh, I'm an idiot. Forgot I had a llama.cpp webserver running with a model loaded. Killed it and it finished very fast.
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#65Earlier quoted context omitted.
FWIW on a 6000 Pro it takes 68 seconds for 10 seconds 480p video, (cold) same demo workflow as you used. Set "megapixels" to 2.0 (1920x1088) and same video seems to take 5+ minutes, not sure if everything is right/correct at the moment. As far as I can tell, the current ComfyUI nodes don't even do compilation, and I haven't looked into what attention mechanism they're using, but I'm sure with time these durations wil…
People on reddit have definitely pointed out that sageattention will speed up the renders. And it's literally the first day. Someone will make a distilled 4-8 step LoRA and we're off to the races. Edit: did a couple of 10 second long 864/480 i2v videos on my RTX Pro 6000: sageattention bumps them up 33%, that is to say, 140.89 seconds without sageattention becomes 105.69 with sageattention on (if using the KJ Sageatt…
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#66The mouse render is surprisingly good. Several of those clips stood out to be a pretty big leap in terms of current SOTA models. The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering. I feel like for a good while now we'll transition into a process that uses traditional "close-u…
This is still about a year and a half behind Seedance 2.0/ Seedance 2.5 But it represents a coming price pressure that will face the leading foundation models. Open source will prevent runaway costs. Moreover, it prevents the hair-trigger platform safety checkers from shutting down creative work. Video models are notoriously bad at shutting down a huge number of requests. Creatives will prefer to work on cloud or pri…
According to one user preference leaderboard, MiniMax H3 is already ahead of Seedance 2.0 based on thousands of A/B votes: https://artificialanalysis.ai/video/leaderboard/text-to-vide...
I haven't seen any user preference comparisons between Seedance 2.5 and MiniMax H3. As an upper bound, H3 cannot be more than 6 months behind Seedance 2.5 since H3 is already ahead of where Seedance was 6 months ago.
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#67Earlier quoted context omitted.
I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics. Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?
> work with types of knowledge that are inherently non-text What are those things exactly? AFAIK, everything we can "know" can be written down, one way or another, even analog circuits. Also, what SOTA LLMs are you referring to? GPTs been handling analog circuits fine for quite some time, I want to say for at least one year? I've been "pair programming" a bunch of working circuits with GPT models since probably GPT 5…
Do you know how to ride a bike?
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#68Has anyone tried running this on Mac device? (e.g. Mac Studio Ultra)
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#69Earlier quoted context omitted.
> work with types of knowledge that are inherently non-text What are those things exactly? AFAIK, everything we can "know" can be written down, one way or another, even analog circuits. Also, what SOTA LLMs are you referring to? GPTs been handling analog circuits fine for quite some time, I want to say for at least one year? I've been "pair programming" a bunch of working circuits with GPT models since probably GPT 5…
> everything we can "know" can be written down Do you know how to ride a bike?
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#70The mouse render is surprisingly good. Several of those clips stood out to be a pretty big leap in terms of current SOTA models. The only one that looks "off" is the beverage ad video during the can opening clip, it still has that "AI smoothening" effect. Good thing this can be done pretty well using traditional rendering. I feel like for a good while now we'll transition into a process that uses traditional "close-u…
right before that, there's the part with the person hiking up the dish and 'breathing', and the clouds of water vapor coming out of their mouth don't line up with their breaths devs pls fix
Really poor. In the last shot, looks more like smoke.
The fact the director even included this item in the reel speaks volumes.