The only thing I'm not sure about is if this model supports thinking or not.
LFM2-24B-A2B: Scaling Up the LFM2 Architecture
11–15 of 15 posts
Re: LFM2-24B-A2B: Scaling Up the LFM2 Architecture
#12This model is pretty cool if you don't have a GPU - I was able to get I think 20 or 30 tokens per second on CPU (DDR4 ram) alone. (I don't remember if that was with q4 or q8.) Otherwise, if you have a GPU with more than like 4GB of VRAM, there are better models. Gemma4 and Qwen3.6 (or Qwen3.5 if you need the smaller dense models that haven't yet been released for 3.6) are a good place to start.
What are you using for inference? I have a recent intel laptop with 32GB of DDR5 and I am getting at most 25tps with the llama cpp vulkan backend (that is the fastest, I also tried sycl but it is a bit slower)
Re: LFM2-24B-A2B: Scaling Up the LFM2 Architecture
#13- GPQA Diamond: 47.4% vs 84.1% for Qwen
- HLE: 4.4% vs 20.2% for Qwen
- AA Omniscience Accuracy: 6.4% vs 18.9% for Qwen
- AA Hallucination Rate: 30.0% vs 50.3% for Qwen
Re: LFM2-24B-A2B: Scaling Up the LFM2 Architecture
#14This model is pretty cool if you don't have a GPU - I was able to get I think 20 or 30 tokens per second on CPU (DDR4 ram) alone. (I don't remember if that was with q4 or q8.) Otherwise, if you have a GPU with more than like 4GB of VRAM, there are better models. Gemma4 and Qwen3.6 (or Qwen3.5 if you need the smaller dense models that haven't yet been released for 3.6) are a good place to start.
> I was able to get I think 20 or 30 tokens per second on CPU (DDR4 ram) alone What are you using for inference? I have a recent intel laptop with 32GB of DDR5 and I am getting at most 25tps with the llama cpp vulkan backend (that is the fastest, I also tried sycl but it is a bit slower)
Prediction Stats:
Stop Reason: eosFound
Tokens/Second: 21.10
Time to First Token: 1.827s
Prompt Tokens: 42
Predicted Tokens: 187
Total Tokens: 229Re: LFM2-24B-A2B: Scaling Up the LFM2 Architecture
#15Liquid AI have made some awesome models (especially the smaller ones, they are lightning fast). I wish they made a fast small size coder. Did a finetune distill of 0.8B myself and it is in fact working properly, coding like a 30B model, so I know it is possible. Anyway here you have the 24B parameters with 2B active: https://hugston.com/models/lfm2-24b-a2b-q4-k-m