Live data from Hacker News

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

github.com

131–140 of 149 posts

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#132

I know everyone wants to crap all over these setups that are impractical, but this is how progress happens. People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc. Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

>when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

I am 100% sure this won't happen in 10 years time. At least not on a $200 SSD. But I wouldn't be surprised if it ran on $1K to $2K HBF SSD. It will still be better than a $50K Rack.

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#133
post #120

Earlier quoted context omitted.

> but this is how progress happens. This is progress in the same way that a man climbing a tree is making progress toward reaching the moon. This project is essentially the MoE-of-the day, with some platform-related optimizations. > Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips. That won't happen. Projects like this just give the illus…

it’s equivalent take to laugh at first transformers 9 years ago, because they were shit and hardware requirements were immense. IMHO it’s a matter of time until we (consumers) will get the hardware (maybe coupled maybe even more novel techniques). Though I expect it will take another 10 years or more.

The frontier labs will do their best to prevent this from happening. Their financial model won't work if people start running open-weight models on their local hardware.

This is why banning Chinese open-weight AI models is a major policy debate in Washington. The labs can't survive log-term without subsidies, and a ban can act as a subsidy.

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#134
post #85

Earlier quoted context omitted.

Here you have shown yourself that progress slows down and doesnt speed up. 8.9/0.35 = ~25x more performance in 10 years from 2006 to 2016. 104.8/8.9 = ~12x more performance in 10 years from 2016 to 2026. Growth has dropped 50%.

That isn't deceleration, you've just chosen a very selective way to compare. If you use time as a denominator, which is kind of intrinsic when talking about rates of acceleration, you get a very different result. If you graphed .3, 8.9, and 104.4 on the y axis, with years on the x axis, it would be pretty clear that there was in increase in the rate of progress. We went from adding 8 teraflops in a decade, to adding…

>It’s like claiming that a company that goes from making $1 to $1k to $100k to $1mm in a 4 year period has decelerating growth.

Because it is decelerating growth. There is a reason why we use YoY percentage in annual and financial reporting.

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#135

I know everyone wants to crap all over these setups that are impractical, but this is how progress happens. People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc. Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

> but this is how progress happens. This is progress in the same way that a man climbing a tree is making progress toward reaching the moon. This project is essentially the MoE-of-the day, with some platform-related optimizations. > Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips. That won't happen. Projects like this just give the illus…

Technically, this can happen. Politically, this won't be allowed to happen - just like digital media ownership never happened.

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#136
post #132

I know everyone wants to crap all over these setups that are impractical, but this is how progress happens. People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc. Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

>when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips. I am 100% sure this won't happen in 10 years time. At least not on a $200 SSD. But I wouldn't be surprised if it ran on $1K to $2K HBF SSD. It will still be better than a $50K Rack.

[dead]

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#137
post #5

Earlier quoted context omitted.

Isn’t it only writes that kill drives?

Yes for NAND, and I suppose nobody is using mechanical hard drives for this.

I'd like to see someone try it, just to see how incredibly slow and noisy it is.

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#139
post #73

Earlier quoted context omitted.

So every script kiddie gets its Mythos to hack sides and scammers don’t need AI services anymore. That time won’t be as much fun as you think

If that's the natural outcome then the future you're describing is inevitable. Oh well, maybe the last one to leave the internet can turn the lights off.

It’s not a natural outcome it’s a decision

Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

#140
post #10

Earlier quoted context omitted.

Haha 1T on $50k might be a bit hopeful, mate, even at FP8. But I too am hopeful.

We'd bought 4 x $11K Mac Studios at my college and via exo, we had Kimi K2.5 at 30 TPS. Not too wild an idea!

30 tok/s generation? Nice! What was prefill by the way? Also, these are 256 GB RAM studios? Good timing on those!
Post reply on HN