Live data from Hacker News

90T/s on my iPhone llama3.2-1B-fp16

reddit.com

1–2 of 2 posts

Re: 90T/s on my iPhone llama3.2-1B-fp16

#2
I made it! 90 t/s on my iPhone with llama1b fp16

We completely rewrite the inference engine and did some tricks. This is a summarization with llama 3.2 1b float16. So most of the times we do much faster than MLX. lmk in comments if you wanna test the inference and I’ll post a link.