Backstory: Hi All, I recently stumbled upon GGML 4 bit quantized LLMs and the fact that small version of these GGML models (i.e., upto 7b) can run smoothly on a CPU! The GGML 4 bit quantized version of replit’s codeInstruct-3b model, only requires 2GB of RAM! So I quickly tested it and hosted the model on a free HuggingFace Space and it works!
Let me know your thoughts on it!
Show HN: Personal Replit Ghostwriter
personal-replit-ghostwriter.sixftone-mlh.repl.co