Earlier quoted context omitted.
Wow, that’s very interesting. I wish more benchmarks were reported along with the total cost of running that benchmark. Dollars per token is kind of useless for the reasons you mentioned.
Yup, MiniMax M-2.5 is a standout in that aspect. It's $/token is very low, because it reasons forever (fun fact, that's also the reason why it's #1 on OpenRouter, because it simply burns through tokens, and OpenRouter ranking is based on tokens usage)...
Gemini 3.1 Flash-Lite: Built for intelligence at scale
31–34 of 34 posts
Re: Gemini 3.1 Flash-Lite: Built for intelligence at scale
#32What the fuck is this price hike? It was such a nice low end, fast model. Who needs 10 years of reasoning on this model size?? I'm gonna switch some workflows to qwen3.5. There's a lot of tasks that benefit from just having a mildly capable LLM and 2.5 Flash Lite worked out of the box for cheap. Can we get flash lite lite please? Edit: Logan said: "I think open source models like Gemma might be the answer here" Imply…
Re: Gemini 3.1 Flash-Lite: Built for intelligence at scale
#33What the fuck is this price hike? It was such a nice low end, fast model. Who needs 10 years of reasoning on this model size?? I'm gonna switch some workflows to qwen3.5. There's a lot of tasks that benefit from just having a mildly capable LLM and 2.5 Flash Lite worked out of the box for cheap. Can we get flash lite lite please? Edit: Logan said: "I think open source models like Gemma might be the answer here" Imply…
Are there good open models out there that beat gemini 2.5 flash on price? I often run data extraction queries ("here is this article, tell me xyz") with structured output (pydantic) and wasn't aware of any feasible (= supports pydantic) cheap enough soln :/
Re: Gemini 3.1 Flash-Lite: Built for intelligence at scale
#34I'm still clinging to gemini-2.0-flash which I think is free free for API use(?!).