Earlier quoted context omitted.
Haha 1T on $50k might be a bit hopeful, mate, even at FP8. But I too am hopeful.
AMD already demonstrated 1T on strix halo clusters. << $10K at original MSRP.
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
141–149 of 149 posts
Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
#142I know everyone wants to crap all over these setups that are impractical, but this is how progress happens. People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc. Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.
Agree with this. As soon as things get in range for motivated amateurs, progress skyrockets. Has also been the case for things like chess computing; a lot of the progress we made over the last decades there (even before involving neural networks!) happened thanks to software improvements because the problem got so accessible, not just faster hardware. I expect similar trends with AI; I'd expect to get decent, human c…
So true. The reverse is also true - when greedy companies overprice their initial release so that it is out of range of the enthusiastic hobbyist they stall progress and adoption.
This is true for hardware (eg failed Intel Optane, Knights Bridge) and software that does not have a cheap or free basic plan.
Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
#143Earlier quoted context omitted.
I don't understand, if they are only using a subset of the tokens then it's a sparse model. What do you mean by dense?
Nothing to understand. Straight up hallucination. I could have sworn I read that they used a novel architecture where the model is dense but you could select specific layers or something at inference. reread the announcement: just said MoE. Corrected my brain's weights so thanks. https://machinelearning.apple.com/research/introducing-third...
Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
#144I see this at the end of the README > Swiftlet was built in collaboration with Claude Code. Did this really happen (some sort of working with Anthropic or Claude Code team) or is it some kind of requirement when you develop some software with Claude Code (I see the other author is: https://github.com/claude ), or sort of reuse some of its parts? Is it like someone saying "built in collaboration with VS Code" or ".. i…
no Anthropic involvement, I just used Claude Code heavily while building this and putting that in the README felt more honest than not mentioning it. Now that I think about it may be it shuld be "built with claude code" instead of "built in collaboration...". Changed it.
> Swiftlet was built with Claude Code.
This is great and correct. It's a tool. Again, thanks.
PS. I just don't why people downvoted me. I am just someone, after a longish sabbatical/gap, exploring and getting used to the agnatic world, though very slowly :)
Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
#145Earlier quoted context omitted.
it’s equivalent take to laugh at first transformers 9 years ago, because they were shit and hardware requirements were immense. IMHO it’s a matter of time until we (consumers) will get the hardware (maybe coupled maybe even more novel techniques). Though I expect it will take another 10 years or more.
The frontier labs will do their best to prevent this from happening. Their financial model won't work if people start running open-weight models on their local hardware. This is why banning Chinese open-weight AI models is a major policy debate in Washington. The labs can't survive log-term without subsidies, and a ban can act as a subsidy.
Though my bet would be, if USA will go ultra protectionist in this regard - in 10-20 years most world will run Chinese LLMs and hardware for this purpose.
Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
#146At what, 10 tokens per hour? These disk swapping methods all have the same drawbacks - kill your drive early, and slow as hell.
Am I the only one that has no flash lifetime anxiety? I still have drives from more than a decade ago that keep on chugging fine. I remember the time spinning rust was the only option and reliable they weren't. In 30 years of computing I have had more than ten hdds and zero ssds die.
No normal use should wear out a drive in any sensible time in reasonable use and even in unreasonable use they seem to last almost indefinitely. There's more likely to be some other kind of component death before that.
Of course, in staged lab test, it's probably possible to burn one out. I've seen projects do that on memory cards and various *ROM chips but don't recall seeing someone kill SSDs. That would get costly. But I'm almost sure a web search would turn out someone doing that.
I also used to be sure that 90's SCSI HDDs would never really stop running. Just the machines using them became too much work and no utility to keep going. I remember only one of mine that wouldn't start after some years in the storage, but I managed to hammer it back into shape.
Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
#147Earlier quoted context omitted.
Most people are already used to rely on the internet on basically everything. At best, they download a tiny chunk of entertainment from it when they go on a plane, and as soon as they land they immediately abandon that offline chunk. In addition, LLMs, small or large, are highly parallelizable. This means that running on the same machine/GPUs many requests in parallel is significantly more efficient, and the sum of t…
I suspect the economics favor centralized servers, if you only look at the aggregated cost to serve X number of users' tokens. But we could say the same thing about a lot of the computation that iPhones do locally. They could have been much thinner clients, but instead they now have more compute power than desktops had when iPhones launched.
Re: Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
#148Earlier quoted context omitted.
Domestic electricity is free nowadays, certainly for most of the year, as solar plus battery covers your usage for a tiny percentage of the cost of your house.
Only if you don't count the cost of the equipment and installation.
Given the cost of a building is far more than the cost of generating enough power for that building it doesn’t really matter