Hollywood and the film industry on red alert. Too bad. This is AGI.
Unlimited slop "content" to fill decomposed brain shaped vessels, the future is bright!
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
21–30 of 100 posts
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#22> We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality. Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as w…
I remember a paper which was posted on HN a few weeks ago where somebody implemented KAN networks in FPGAs, since those can readily be approximated as LUTs.
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#23Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#24Hollywood and the film industry on red alert. Too bad. This is AGI.
The example video just looks like the highly produced art (TV, commercials, games) other people have created. I find it impossible to believe this wasn't trained on other people's work, and there is no protection for it. Terribly sad. A lack of original thinking is coming.
As with the arts, 99.9% of people can't use these models to express vision, get attention, or achieve distribution.
The game is the same as it has always been. You still need hard work, taste, something important to say, the ability to articulate it, good timing, and luck.
Nothing has changed. We can just build faster.
What this does enable is for more to be created that caters to a wider variety of interests. It disrupts existing structures of capital allocation, production, and distribution and gives new players a chance to reshape the game.
The bar will rise and people will still be running at the same pace on the treadmill. There will be more to see, but less time to see it.
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#25Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.
FWIW on a 5080 16GB it takes 3 minutes for 10 seconds 480p video (the mouse video workflow with length changed from 5 seconds to 10 seconds)
As far as I can tell, the current ComfyUI nodes don't even do compilation, and I haven't looked into what attention mechanism they're using, but I'm sure with time these durations will come down even more.
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#26On one hand: impressive. On the other hands aesthetically it all looks painfully bland and generic.
I remember reading a report where people running AI-model instagram account were using insanely long and detailed prompts about the setting, lighting, makeup, pose, disposition, clothing, etc. about their models. Presumably with some reference image of the face / body to remain consistent across images.
It‘s not clear to me whether a sufficiently detailed prompt can generate actually interesting video with a natural ”texture” (for lack of a better word).
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#27Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#28> We found that the model's modulation weights (~40% of the total parameters) could be pruned and replaced with a functionally equivalent lookup table, dramatically shrinking the memory footprint with no loss in output quality. Is this a common approach to reducing weights with "no loss in output quality", assuming this is true? Seems almost too simple to work. If this is doable, would this be applicable to LLMs as w…
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#29Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.
I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics. Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?
What are those things exactly? AFAIK, everything we can "know" can be written down, one way or another, even analog circuits.
Also, what SOTA LLMs are you referring to? GPTs been handling analog circuits fine for quite some time, I want to say for at least one year? I've been "pair programming" a bunch of working circuits with GPT models since probably GPT 5 or so.
Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
#30Im running this on my 4070ti super (16 gb vram), and it takes 10 minutes for a 10-seconds 480p video. but the results are spectacular.
I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics. Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?