Live data from Hacker News

Small offline large language model – TinyChatEngine from MIT

graphthinking.blogspot.com

11–20 of 25 posts

Re: Small offline large language model – TinyChatEngine from MIT

#11
post #6

Use llama.cpp for quantized model inference. It is simpler (no Docker nor Python required), faster (works well on CPUs), and supports many models. Also there are better models than the one suggested. Mistral for 7B parameters. Yi if you want to go larger and happen to have 32Gb of memory. Mixtral MoE is the best but requires too much memory right now for most users.

Python is only used in the toolchain, the inference engine is entirely C/C++.

Re: Small offline large language model – TinyChatEngine from MIT

#12
I tried this and installation was easy on macOS 10.14.6 (once I updated Clang correctly).

Performance on my relatively old i5-8600 CPU running 6 cores at 3.10GHz with 32GB of memory gives me about 150-250 ms per token on the default model, which is perfectly usable.

Re: Small offline large language model – TinyChatEngine from MIT

#13
post #3

I’m a tad confused > TinyChatEngine provides an off-line open-source large language model (LLM) that has been reduced in size. But then they download the models from huggingface. I don’t understand how these are smaller? Or do they modify them locally?

https://github.com/mit-han-lab/TinyChatEngine Turns out the original source is actually somewhat informative. Including telling you how much hardware do you need. This blog post looks like your typical note you leave for yourself to annotate a bit of your shell history.

Your assessment is exactly correct -- the blog post is my note-to-self about getting the repo to work. My "added value" in the post is a Dockerfile for ease of installation.

Re: Small offline large language model – TinyChatEngine from MIT

#14
post #6

Use llama.cpp for quantized model inference. It is simpler (no Docker nor Python required), faster (works well on CPUs), and supports many models. Also there are better models than the one suggested. Mistral for 7B parameters. Yi if you want to go larger and happen to have 32Gb of memory. Mixtral MoE is the best but requires too much memory right now for most users.

Thanks for the suggestion. I'm new to running LLMs so I'll take a look at your suggestion [0]. My ~10 year old MacBook Air has 4GB of RAM, so I'm primarily interested in smaller LLMs.

[0] https://github.com/ggerganov/llama.cpp

Re: Small offline large language model – TinyChatEngine from MIT

#16

I have used them and I can say it's pretty decent overall. I personally plan to use tinyengineon iot devices which is for even smaller iot microcontroller devices.

May I ask what your use case is? I've found LLMs are pretty good at parsing unstructured data into JSON, with minimal hallucinations.

Is there a tutorial on how to do something like that? It sounds damn useful.

Re: Small offline large language model – TinyChatEngine from MIT

#17
post #6

Use llama.cpp for quantized model inference. It is simpler (no Docker nor Python required), faster (works well on CPUs), and supports many models. Also there are better models than the one suggested. Mistral for 7B parameters. Yi if you want to go larger and happen to have 32Gb of memory. Mixtral MoE is the best but requires too much memory right now for most users.

Thanks for the suggestion. I'm new to running LLMs so I'll take a look at your suggestion [0]. My ~10 year old MacBook Air has 4GB of RAM, so I'm primarily interested in smaller LLMs. [0] https://github.com/ggerganov/llama.cpp

You don't necessarily need to fit the model all in memory – llama.cpp supports mmaping the model directly from disk in some cases. Naturally inference speed will be affected.

Re: Small offline large language model – TinyChatEngine from MIT

#19
post #6

Use llama.cpp for quantized model inference. It is simpler (no Docker nor Python required), faster (works well on CPUs), and supports many models. Also there are better models than the one suggested. Mistral for 7B parameters. Yi if you want to go larger and happen to have 32Gb of memory. Mixtral MoE is the best but requires too much memory right now for most users.

I'm curious, what do you use these small LLMs for, like can you give some examples of (not too) personal uses cases from the past month?

Re: Small offline large language model – TinyChatEngine from MIT

#20

I have used them and I can say it's pretty decent overall. I personally plan to use tinyengineon iot devices which is for even smaller iot microcontroller devices.

May I ask what your use case is? I've found LLMs are pretty good at parsing unstructured data into JSON, with minimal hallucinations.

I’m also curious.
Post reply on HN