Will it fit on ESP32??
Running Kimi K3 on a M1 Max
11–20 of 97 posts
Re: Running Kimi K3 on a M1 Max
#120.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?
Also its answering the question of what gonna happen if you wake up tomorrow and datacenters are gone. Or internets are gone.
Some people on our globe live in countries with no internet whatsoever. Of course most of them dont have Macbook with 64GB RAM either, but it's much much easier to get than internet connection or rack of GB200.
SOTA LLMs are efficiently compression of all the knowkedge humanity has built. Having ability to run it at home to extract said knowledge is important no matter the speed.
Re: Running Kimi K3 on a M1 Max
#130.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?
I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.
Re: Running Kimi K3 on a M1 Max
#14Re: Running Kimi K3 on a M1 Max
#15Earlier quoted context omitted.
I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.
Love the email idea
It reminds me of when Willow Garage chose to name their bot the TurtleBot, because if they named it anything else, people would think it was fast and capable. But when they called it Turtle Bot, people just kind of liked it and were satisfied with what it did.
At the level of Kimi 3, I probably can code only about 1,000 good tokens per day, too. (thankfully coding isn't my job)
Re: Running Kimi K3 on a M1 Max
#16Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).
Re: Running Kimi K3 on a M1 Max
#170.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?
I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.
Kimi Pen Pal. Bring back lettets and postcards. Do OCR, and use one of those 3D printer-like pen plotters write the model output as a letter.
Challenge would be automating the opening and OCR preparation, and the folding and mailing of the return letter. But given it's done commercially it should be possible.
Re: Running Kimi K3 on a M1 Max
#18Says it requires a 2TB disk? Must it be internal NVMe?
Re: Running Kimi K3 on a M1 Max
#19Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).
Reads are not generally life-limiting for flash. (Well, no more so than power-on time in general. You still have aging mechanisms like electromigration, but these are orders of magnitude slower than write-induced damage.)
Re: Running Kimi K3 on a M1 Max
#20Will it fit on ESP32??