H3-metal β Native MiniMax-H3 inference for Apple Silicon
199 points - today at 1:22 AM
SourceMeleagris
today at 3:12 AM
I've been using MiniMax H3 on my M5 Pro 64GB MacBook Pro through ComfyUI. It works extremely well.
I had to modify the default ComfyUI workflows to use a GGUF quant (city96's ComfyUI-GGUF custom node, UnetLoaderGGUF in place of the stock loader) [0].
I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest.
The main issue is speed, a ~9-second 480x864 clip at 20 steps takes me a bit over an hour. So this will be cool to try for the speed up alone.
There's a lot of great information and workflows available to follow on the r/StableDiffusion subreddit.
[0] https://huggingface.co/Abiray/MiniMax-H3-GGUF/tree/main/unet
alexgoodhart
today at 5:44 AM
I wonder how much faster your m5 pro is compared to my M1 Max @ 64gb
linzhangrun
today at 6:39 AM
On my 128GB M4 Max Mac Studio, generating a 15s 480p video with MiniMax H3 in ComfyUI takes an hour and a half.
Put Codex to work on deploying it now, hoping the speed can improve quite a lot :-) Thanks anyway
sscarduzio
today at 6:48 AM
Please keep up posted about the results!
In the AMA Minimax said that H3 could support sparse attention, that would be a huge speedup! I wonder if there are any news on that. H3 is very cool. EDIT: testing a --sparse-attention optional mode based on what they said in the Reddit post.
This is where the DGX spark makes up a bit of the ground it loses on llm work, diffusion and cuda go together like peanut butter and jelly.
TechSquidTV
today at 2:42 AM
This still requires 128Gb of memory, right? Me and my lowly 96Gb, like a commoner; missing out on the fun.
thehamkercat
today at 2:46 AM
From README:
> On the 128 GB M5 Max, clean end-to-end image+audio and embedded-video+audio renders completed in 74.58 and 76.99 seconds respectively, each with about a 40.1 GB peak physical footprint and zero swaps.
Looks like it uses 40GB? So your 96GB mac setup should work fine i guess (Model itself is 33B)
This repo looks neat, but I hope they add some more clear benchmarks because that time (74.58s) is pretty meaningless given that the it/s (and total time) is highly dependent on mode (T2V vs I2V vs REF2V), resolution (0.4, 0.6mp, etc), duration (5-15 seconds), etc.
c0rruptbytes
today at 3:22 AM
wow antirez does not sleep
Wow, had to check some of his other repos; his the one behind dump1090
matheusmoreira
today at 6:54 AM
He's the one who wrote Kilo too!
Understatement of the year :-D
when you have enough money to not have to worry about anything, you can go back to your hobbies. in this case, his hobby is programming.
Being a world class talent is independent of financial situation.
talent w/o financial stability is a battery w/o circuit.
stressback
today at 4:31 AM
"enough money not to worry about anything" haha
I'd love to know what the alternatives are and how this is better
How similar are Jeff Dean and Salvatore Sanfilippo?
onionisafruit
today at 2:59 AM
My favorite Jeff Dean fact is that heβs also antirez. Which reminds me of my favorite Salvatore Sanfilippo fact. Heβs also Jeff Dean
I'm totally following this
Is that what the identity function is?
It's why javascript had to add the triple equals check...
muragekibicho
today at 6:16 AM
2 is not enough. 3 verifies the Dean-Sanfilippo correspondence.
robotresearcher
today at 6:19 AM
People are really good at stuff.
I noticed on a bar TV the other day that some of the Chromecast screensaver landscape photo credits were to Peter Norvig. They were really lovely pictures.
songhonglei1985
today at 4:49 AM
[dead]