Live data from Hacker News

Inkling: Our Open-Weights Model

thinkingmachines.ai

101–110 of 324 posts

Re: Inkling: Our Open-Weights Model

#101
post #34

Raised 2 billion dollars at a 12 billion valuation and debuts at 41 on the Artificial Analysis Intelligence Index, while KIMI and DeepSeek will release Fable-class models this week. What a joke.

Anyone would think these investors are making a bet they can improve using that cash.

Re: Inkling: Our Open-Weights Model

#102
Very preliminary testing so far, but there is something here, far beyond what the benchmarks suggest. Only ever saw such outperformance of public evals vs my private ones with Anthropic models and while it is far to early to make any judgement at this stage, this model will take up a lot of mine time in the coming weeks by the look of things. Only ever viewed Moonshot AIs models as something I'd be able to live with open-weight-wise (Z.AIs output simply does not perform as well in my task set), but this has the potential to be the second. If Mistral came out with something like this, I suspect every Europhile (me included) would never stop talking about it.

Re: Inkling: Our Open-Weights Model

#103
post #36

Earlier quoted context omitted.

Just serving the model over API seems like a natural fit and is what many of them are doing. So simply being the cloud provider for your own open weight model can be a source of revenue

But so can everyone else. What’s the moat for spending all those billions. I understand the Chinese angle, they need to undermine American models as a matter of statecraft, but what is the business model here? It just seems like VC charity.

There are no moats. LLM's are a commodity. The point in spending all of the billions is to have strong domestic open-weight models.

One of the worst case scenarios regarding LLM's is monopoly control, so these billionaires know they need to invest in competition.

Re: Inkling: Our Open-Weights Model

#104
What strikes me the most is just how many different tasks are involved in modern model design. It used to be the case that you come up with a new loss function, slight architecture changes, etc., run your train and eval loop, and publish the artifacts.

Now, there’s so much work to do just to keep up. It’s the ultimate red queen race. All of the 500 steps involved, each of which is its own little optimization loop, is sort of awe inspiring.

But obviously this inverts the previous rules that small teams run faster than big teams. AI requires a big team. It’s only once the team pushes past the 1000s that organizational inertia seems to become an issue. Because until then, there’s way too many pieces for even a dozen super stars.

Re: Inkling: Our Open-Weights Model

#105
post #2

America needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice. Thinking Machines might be it.

It’s what Meta was supposed to do but Llama fell of the wagon.

There’s also Prism

Re: Inkling: Our Open-Weights Model

#106
I think we’re going to start seeing more OSS models that perform especially well on certain tasks instead of trying to be generalists like the frontier models. That’s a winning formula because if you’re building an app on a model it often has a specific set of use cases

Re: Inkling: Our Open-Weights Model

#107
post #2

America needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice. Thinking Machines might be it.

Also the fact that China is building solar power like crazy: that makes it fantastically more well spirited an endeavor to wish well.

Re: Inkling: Our Open-Weights Model

#108
post #7
post #2

America needs its own DeepSeek or Z.ai, a lot of people (myself included) root for open chinese models to win because they have no other choice. Thinking Machines might be it.

It could be but there are a host of companies going after open weights models: Arcee, Reflection, Llama (TBD on Meta's focus on closed-source versus open-source), etc. That said, the fine-tuning API + open weight model at least is a semblance of a viable business that could work so I will be curious about it. I'm not sure the synergy is fully there (why is someone with an open weights model privelaged to fine-tune it…

Do any of these even have match a year old Deepseek 3.1?
Post reply on HN