Live data from Hacker News

H3-metal – Native MiniMax-H3 inference for Apple Silicon

github.com

61–70 of 108 posts

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#62

Earlier quoted context omitted.

It shouldn't? Unless you're using BF16 for all weights (I'm using NVFP4 for the text encoder, otherwise everything BF16 (and audio F32)) you'll fit it all within 96GB VRAM, bugs non-with-standing :) I've been fitting this within 96GB VRAM without issues.

*notwithstanding Anyway, good input!

I'll blame it on other book authors! :) https://en.wiktionary.org/wiki/nonwithstanding

> This misconstruction is very common, included in print publications spanning several centuries. It might be considered an alternative spelling, albeit still a mistaken usage.

Thanks though, I never actually knew so was helpful :)

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#64

Earlier quoted context omitted.

*notwithstanding Anyway, good input!

I'll blame it on other book authors! :) https://en.wiktionary.org/wiki/nonwithstanding > This misconstruction is very common, included in print publications spanning several centuries. It might be considered an alternative spelling, albeit still a mistaken usage. Thanks though, I never actually knew so was helpful :)

[dead]

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#65
post #29

Earlier quoted context omitted.

Of course it isn’t. If you can’t afford to eat you can’t achieve any potential you might have. Financial stability is a gamechanger for everyone.

That's stupid, if you're truly talented you'll solve the financial stuff in order to pursue whatever you want to do - if you don't then that's on you.

You should read Outliers by Malcolm Gladwell

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#66
post #39

Earlier quoted context omitted.

There will Reddit subs for it though couldn’t tell you which off top of my head I’d personally steer clear of messaging platforms for this - who knows what one might stumble into there

> I’d personally steer clear of messaging platforms for this - who knows what one might stumble into there Personally I have no interest, but sometime browse stuff out of curiosity. But this got more of my curiosity, what kind of "stuff" are you implying they might stumble upon on the open, public internet? Sure, some NSFW, horror and otherwise weird stuff is there, especially around AI generation, but hardly somethi…

I do not know and very much plan to keep it that way

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#68

I've been using MiniMax H3 on my M5 Pro 64GB MacBook Pro through ComfyUI. It works extremely well. I had to modify the default ComfyUI workflows to use a GGUF quant (city96's ComfyUI-GGUF custom node, UnetLoaderGGUF in place of the stock loader) [0]. I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest. The main issue is spee…

> a ~9-second 480x864 clip at 20 steps takes me a bit over an hour

that's rough. for comparison, i tried the exact same parameters on my 5090 RTX and it took 2 minutes to generate.

i believe diffusion models are primarily compute bound so the macs aren't really the ideal hardware for this kind of stuff

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#69

Earlier quoted context omitted.

*notwithstanding Anyway, good input!

I'll blame it on other book authors! :) https://en.wiktionary.org/wiki/nonwithstanding > This misconstruction is very common, included in print publications spanning several centuries. It might be considered an alternative spelling, albeit still a mistaken usage. Thanks though, I never actually knew so was helpful :)

[deleted]

Re: H3-metal – Native MiniMax-H3 inference for Apple Silicon

#70

Earlier quoted context omitted.

you should have a look at https://github.com/deepbeepmeep/Wan2GP which is the goto tool for "gpu poor", although as people below already pointed out you should be fine with comfyui's standard setup aswell

First, I think they're not even talking about GPUs, this is macOS hardware so unified memory. Secondly, if they were talking about GPUs, then 96GB VRAM is hardly what people refer to when they say "gpu poor".

1) It would still run on the Mac's GPU.

2) Since it's unified memory, you won't have 96GB available.

3) I offered a solution that is usually recommended to the "gpu poor", if he's concerned with how much memory he would need.

4) I stated, that people already pointed out how he should be fine and that "gpu poor" doesn't apply to him.

5) "gpu poor" depends on what model you are trying to use. If you want to run Kimi or GLM you are still "gpu poor" even if you have an RTX Pro 6000 with 96GB of VRAM.

Post reply on HN