Earlier quoted context omitted.
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes. Huge pric…
> Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really re…
so I have a a few decades worth of creative coding projects, many personal notes + writing, pictures, video clips, renders, projects, backups of the previous laptop's home folder that contains the laptop before it in the ~/dump folder and so on, spread out over several harddisks, laptops and storage media
it's a mess that I've wanted to organize for years, but I can't seem to manage it
but I might be able to sketch out a plan or task for a local multi-modal LLM to slowly chug through all the files (read-only), catalog and index them, copy to a clean external SSD, de-dupe and tidy everything up
I mean, some of this data is already highly personal, and there might be stuff I forgot was there
... is it an "extreme data privacy requirement" to not want to share running an LLM for this task, and wanna do it locally?