Live data from Hacker News

Inkling: Our Open-Weights Model

thinkingmachines.ai

21–30 of 324 posts

Re: Inkling: Our Open-Weights Model

#21
post #4

If it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model? Maybe for the multi modal?

> If it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model? The benchmarks never tell the full story. Some of the open weights models have been benchmaxxed for a while. Their utility on real work can be different than the benchmark number. The multimodal input is also a big deal. Having vision input is really helpful for a lot of tasks.

I second that. Gemini 3.5 Flash rocks the benchmark charts but is terrible as an agent. Horrible instruction adherence and makes WAY too many tool calls

Re: Inkling: Our Open-Weights Model

#24
post #20

Earlier quoted context omitted.

> If it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model? The benchmarks never tell the full story. Some of the open weights models have been benchmaxxed for a while. Their utility on real work can be different than the benchmark number. The multimodal input is also a big deal. Having vision input is really helpful for a lot of tasks.

Then why are they publishing the benchmarks which makes them look worse than GLM 5.2?

being close is still impressive, especially for their first (released) model

gives me hope that the training moat is even smaller than we thought

Re: Inkling: Our Open-Weights Model

#29

For the most part it’s better than Nemotron, worse than GLM. This makes it the best American open weights model from what I can tell?

It's nearly double the size of Nemotron 3 Ultra, so I'd expect it to be considerably better, although the active parameter count seems to be a touch lower at 41B vs 55B

Re: Inkling: Our Open-Weights Model

#30

Earlier quoted context omitted.

> If it's ~30% bigger and not as good as GLM 5.2, why would I tinker with this model? The benchmarks never tell the full story. Some of the open weights models have been benchmaxxed for a while. Their utility on real work can be different than the benchmark number. The multimodal input is also a big deal. Having vision input is really helpful for a lot of tasks.

I second that. Gemini 3.5 Flash rocks the benchmark charts but is terrible as an agent. Horrible instruction adherence and makes WAY too many tool calls

which cheap models have you found work best as agents?
Post reply on HN