90T/s on my iPhone llama3.2-1B-fp16
reddit.com
90T/s on my iPhone llama3.2-1B-fp16
1–2 of 2 posts
Re: 90T/s on my iPhone llama3.2-1B-fp16
#2I made it! 90 t/s on my iPhone with llama1b fp16
We completely rewrite the inference engine and did some tricks. This is a summarization with llama 3.2 1b float16. So most of the times we do much faster than MLX. lmk in comments if you wanna test the inference and I’ll post a link.