Live data from Hacker News

Running Kimi K3 on a M1 Max

github.com

21–30 of 97 posts

Re: Running Kimi K3 on a M1 Max

#21
post #7
post #3

0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?

I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.

Email would indeed be fitting for K3 running on a M1 Mac, as it'd take days/weeks to receive a response, which matches with my real-world emailing experience pretty well.

Re: Running Kimi K3 on a M1 Max

#24
post #3

0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?

So you subscribe to the belief we won't in future find mentalism in other galaxies or solar systems which operate on mechanisms we don't understand and think v e r y s l o w w w w w w w l y ?

(note. I am not a believer in AGI)

"useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.

Re: Running Kimi K3 on a M1 Max

#26

> ~60–76 s/token I don't know if I'd call this "running"

I commented similarly below, but as a terrible programmer, I probably perform about 1 minute per token too (at Kimi 3 level). It puts into context how I think about intelligence

Re: Running Kimi K3 on a M1 Max

#28
post #26

> ~60–76 s/token I don't know if I'd call this "running"

I commented similarly below, but as a terrible programmer, I probably perform about 1 minute per token too (at Kimi 3 level). It puts into context how I think about intelligence

Maybe in terms of code produced, but one token is only a fragment of a thought for an LLM.

It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.

Post reply on HN