Live data from Hacker News

Running Kimi K3 on a M1 Max

github.com

31–40 of 97 posts

Re: Running Kimi K3 on a M1 Max

#32
post #24
post #3

0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?

So you subscribe to the belief we won't in future find mentalism in other galaxies or solar systems which operate on mechanisms we don't understand and think v e r y s l o w w w w w w w l y ? (note. I am not a believer in AGI) "useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.

Are there any well thought through stories about what this would look like? For example, I'm thinking about like nutrient flow, decision making, energy input, gravitational force, things like that seem to govern the value and speed of intelligence.

Re: Running Kimi K3 on a M1 Max

#33

Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).

No problem at all to read data over and over. In fact, LLM weights are a great candidate for low-quality flash that can't handle a lot of write cycles, and you want a large amount of storage cheaply...

Re: Running Kimi K3 on a M1 Max

#35
post #26

Earlier quoted context omitted.

I commented similarly below, but as a terrible programmer, I probably perform about 1 minute per token too (at Kimi 3 level). It puts into context how I think about intelligence

Maybe in terms of code produced, but one token is only a fragment of a thought for an LLM. It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.

Oh, is that how it works? So, when somebody says a model is running at X tokens per second, it means that the thinking process is running at that, and output tokens are much lower then? Thanks to the explanation.

Re: Running Kimi K3 on a M1 Max

#36
Now set it up with an agent and a permanent `/goal` to say it cannot stop until it has solved for speed, then leave it on and livestream so we can all see when it becomes exponential. Could have the Eternal Jukebox playing in the background!

Re: Running Kimi K3 on a M1 Max

#38
post #4

SSD streaming on an M5 Max 128GB: https://x.com/antirez/status/2082136334160818528 Soon decent speed across two Mac Studios with 512GB of RAM.

Cool stuff. Do you have a the hardware and a way to bridge the compute? Or just hopeful?

Re: Running Kimi K3 on a M1 Max

#40
post #3

0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?

16 tokens / s is not nothing.

It's the other way around due to poor framing, it would be much easier to compare if you [the repo] said 0.02 tps.
Post reply on HN