Live data from Hacker News

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

blog.comfy.org

71–80 of 100 posts

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#71
post #32

Earlier quoted context omitted.

This is still about a year and a half behind Seedance 2.0/ Seedance 2.5 But it represents a coming price pressure that will face the leading foundation models. Open source will prevent runaway costs. Moreover, it prevents the hair-trigger platform safety checkers from shutting down creative work. Video models are notoriously bad at shutting down a huge number of requests. Creatives will prefer to work on cloud or pri…

A year and a half behind Seedance 2.0? That is a bold claim that needs evidence. According to one user preference leaderboard, MiniMax H3 is already ahead of Seedance 2.0 based on thousands of A/B votes: https://artificialanalysis.ai/video/leaderboard/text-to-vide... I haven't seen any user preference comparisons between Seedance 2.5 and MiniMax H3. As an upper bound, H3 cannot be more than 6 months behind Seedance 2…

> A year and a half behind Seedance 2.0? That is a bold claim that needs evidence.

SOTA a year ago was Kling 2.5, and H3 does not look or perform at that level.

The Artificial Analysis rankings are whack. They rank Omni first, which is incredulously wrong. Google's models broadly suck, and there they all are - right at the top.

Artificial Analysis has notoriously ranked models such as Grok Imagine highly and continues to rank Happy Horse as a good model despite the model being absolute garbage.

Could it be because they are subject to broad based statistical attacks? It's easy to encode information about the origin of media in either its metadata or output frames. Or maybe there's simply no overlap between creatives and people who click on ELO scores.

I've spent thousands upon thousands of dollars generating video. I will stand by the claim that nothing touches Seedance 2.0 / 2.5.

It's good that we're getting better open weights. It puts price pressure on the foundation model companies. But these open models are not a substitute for Kling or Seedance yet. Not even close.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#72
post #4
post #3

Hollywood and the film industry on red alert. Too bad. This is AGI.

The example video just looks like the highly produced art (TV, commercials, games) other people have created. I find it impossible to believe this wasn't trained on other people's work, and there is no protection for it. Terribly sad. A lack of original thinking is coming.

It is coming up with Seinfeld episodes even with voice overs. Unless it is trained to do so, don't think that would happen. This model is just around the corner from getting banned due to this copyright issue. (which i hope not)

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#73
post #26

On one hand: impressive. On the other hands aesthetically it all looks painfully bland and generic.

I agree. But is that the model or the prompt? I remember reading a report where people running AI-model instagram account were using insanely long and detailed prompts about the setting, lighting, makeup, pose, disposition, clothing, etc. about their models. Presumably with some reference image of the face / body to remain consistent across images. It‘s not clear to me whether a sufficiently detailed prompt can gener…

> But is that the model or the prompt?

All looks...

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#74
post #67

Earlier quoted context omitted.

> work with types of knowledge that are inherently non-text What are those things exactly? AFAIK, everything we can "know" can be written down, one way or another, even analog circuits. Also, what SOTA LLMs are you referring to? GPTs been handling analog circuits fine for quite some time, I want to say for at least one year? I've been "pair programming" a bunch of working circuits with GPT models since probably GPT 5…

> everything we can "know" can be written down Do you know how to ride a bike?

I think that's different though, that's "doing" rather than just "knowing". You can ride a bike without knowing how it works, and obviously vice-versa too. I don't see circuit diagrams as "doing" though, but the soldering part of building circuits definitely is that way though, you can't just read about it and excel first time you pick up an iron, you have to practice and understand it with your body, like bicycling.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#75
post #36
post #18

Earlier quoted context omitted.

I am particularly curious how multimodal models will work with types of knowledge that are inherently non-text. For example, SOTA LLMs really suck at electronics, especially analog electronics. Is MiniMax H3 capable of logical / technical reasoning, or is it purely art oriented?

Suck at what aspect of (analog) electronics specifically? Not contradicting the claim, just want to understand it. I have not tested yet, but I suspect that LLMs with a harness that can execute code can do SPICE simulations rather ok these days? I have seen MCPs for measurement equipment also, maybe they can even close the physical loop?

> I have seen MCPs for measurement equipment also, maybe they can even close the physical loop?

I'm working on this but for various reasons can't have my physical lab up at the moment. But yes there are many options for connected test equipment that could rather trivially be interacted with via LLM or pretty easy to write libraries.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#76
post #49

Earlier quoted context omitted.

I've had great results on electronics with Claude and Gemini.

what sorts of things have you designed with it?

Reverse-engineering my house's intercom system.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#78
post #67

Earlier quoted context omitted.

> everything we can "know" can be written down Do you know how to ride a bike?

I think that's different though, that's "doing" rather than just "knowing". You can ride a bike without knowing how it works, and obviously vice-versa too. I don't see circuit diagrams as "doing" though, but the soldering part of building circuits definitely is that way though, you can't just read about it and excel first time you pick up an iron, you have to practice and understand it with your body, like bicycling.

> that's "doing" rather than just "knowing"

Yet I know it even when I am not doing it.

I bet you have dozens of skills you cannot represent in words.

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#79
I dug up a few old parody ideas I’d had back in high school and threw them at MiniMax M3 on my RTX. There’s definitely still a lot of jank once you move away from fairly normal scenarios. The moment you start to veer into weirder concepts, things tend to break down a bit especially in the game show where someone is strapped to a wheel and being spun.

Still tho, I was actually shocked by how well the text-to-video turned out overall, and how fast it ran. A 10-second, half-megapixel video gen took only a few minutes which is kinda crazy especially thinking back early WAN days.

Video demos:

https://imgpb.com/rllwg

Re: MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

#80

I dug up a few old parody ideas I’d had back in high school and threw them at MiniMax M3 on my RTX. There’s definitely still a lot of jank once you move away from fairly normal scenarios. The moment you start to veer into weirder concepts, things tend to break down a bit especially in the game show where someone is strapped to a wheel and being spun. Still tho, I was actually shocked by how well the text-to-video tur…

Hilarious well done
Post reply on HN