This is a very flashy page that's glossing over some pretty boring things. - This is a benchmark for "home security" workflows. I.e., extremely simple tasks that even open weight models from a year ago could handle. - They're only comparing recent Qwen models to SOTA. Recent Qwen models are actually significantly slower than older Qwen models, and other open weight model families. - Specific tasks do better with spec…
You are very correct, I just have 2 days of the MBP PRO 64GB on hands, so the test is just covering LLM part -- the logic handling.
For VLM, LFM is the best, even 450M works, I'll update soon :) Thanks again for your deep understanding of LLM/VLM domain and your suggestion.