Live data from Hacker News

H3-metal – Native MiniMax-H3 inference for Apple Silicon

github.com

81–90 of 108 posts

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#82
post #29

Earlier quoted context omitted.

Of course it isn’t. If you can’t afford to eat you can’t achieve any potential you might have. Financial stability is a gamechanger for everyone.

That's stupid, if you're truly talented you'll solve the financial stuff in order to pursue whatever you want to do - if you don't then that's on you.

The majority of the population doesn't even have access to a functioning computer. So yeah maybe somebody truly talented can figure their way out of that hole after a few years but that's where alot of people are starting from.

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#83

I'd love to know what the alternatives are and how this is better

This will be a little faster right now on an M4 or M5 because it's optimized for Apple Silicon. Assuming this model is still state of the art in six months, which might not be a surprise given how long other video models have taken, this should be much, much faster with the M7 chip.

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#84

On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half. Put Codex to work on deploying it now, hoping the speed can improve quite a lot :-) Thanks anyway

> On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half. That's crazy, a RTX Pro 6000 does that in in 2-3 minutes (give or take, depending on your exact settings). LLMs don't make the difference between standalone GPU vs unified memory + CPU so obvious as diffusion models seems to do.

An RTX6000 is a completely different class of hardware.

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#85
post #71

I've been using MiniMax H3 on my M5 Pro 64GB MacBook Pro through ComfyUI. It works extremely well. I had to modify the default ComfyUI workflows to use a GGUF quant (city96's ComfyUI-GGUF custom node, UnetLoaderGGUF in place of the stock loader) [0]. I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest. The main issue is spee…

GGUF is outdated in the latest versions of Comfy-UI. If you want a good balance of size, speed and quality you should use the int8_convrot model from the official Comfy Org Repo https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffus...

So I did test this, and it doesn't work because the quantized layers need torch._int_mm, which PyTorch's MPS backend doesn't implement. It just throws NotImplementedError.

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#86

On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half. Put Codex to work on deploying it now, hoping the speed can improve quite a lot :-) Thanks anyway

For gods sakes, Apple let people have run other GPUs instead of these pissweak 2012 class mobile GPUs

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#87
post #54

Alright I’ve been afraid to ask but have been having trouble finding What are some adult entertainment workflows in comfyui, I need best loras, best prompts to start with and the communities, are they on telegram or something?

H3 is quite uncensored, but was not trained on p0rn, so it has no anatomy clues needed to generate that kind of stuff. For softer adult content it is reported to be fine on Reddit.

loras have been fixing that for years

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#89

Earlier quoted context omitted.

> I’d personally steer clear of messaging platforms for this - who knows what one might stumble into there Personally I have no interest, but sometime browse stuff out of curiosity. But this got more of my curiosity, what kind of "stuff" are you implying they might stumble upon on the open, public internet? Sure, some NSFW, horror and otherwise weird stuff is there, especially around AI generation, but hardly somethi…

One thing I read on this topic on reddit is never EVER use the word "girl" when prompting H3. So CSAM probably.

Did you try this yourself? Of course you wouldn't, because not wanting to produce SCAM sorry I meant CSAM.

And no, including the word "girl" in H3 does not lead to CSAM in any way, shape or form, but it's a great example how FUD quickly spreads.

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#90

Earlier quoted context omitted.

> On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half. That's crazy, a RTX Pro 6000 does that in in 2-3 minutes (give or take, depending on your exact settings). LLMs don't make the difference between standalone GPU vs unified memory + CPU so obvious as diffusion models seems to do.

An RTX6000 is a completely different class of hardware.

Really? No wonder I keep trying to type on it like a laptop but it doesn't work and doesn't even have a display!
Post reply on HN