Kubernetes is a system/infrastructure orchestration tool. I completely fail to see how it is comparable to open weight neural nets. In either application or function. I'm sorry to do that hn comment thing where we all just race to contradict or talk in opposition of whatever was said before. I'm aware. But really guys, was this article really not just a miss?
Open-weight AI is having its Kubernetes moment
311–320 of 346 posts
Re: Open-weight AI is having its Kubernetes moment
#312Open-weight and OSS are wildly different and the article makes a poor comparison. What's the incentive for the Chinese labs to continue releasing weights 5 years from now? It's not a stable equilibrium and cannot last. - The lab spending large sums on research and training does not get the inference revenue to fund those efforts. - Unlike OSS where a single volunteer can keep a project going, training costs run into…
China has no end of money to support these companies. The reason this equilibrium is unstable is that the autonomous agentic coding aspect of the models has been so successfully improved that it will soon be a threat to China state security.
Re: Open-weight AI is having its Kubernetes moment
#313One of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything. So what open weight models do is at least provide a baseli…
> if you really want Kimi K2 instead of K3 you can still use it. I think this is a very important aspect, especially after the huge GPT-4o backlash when GTP-5 came out. Each model has certain quirks, and areas where the previous model might be better for some use cases than the latest and greatest, and the labs so far seem to have no desire to offer some kind of "LTS" release.
Re: Open-weight AI is having its Kubernetes moment
#314Re: Open-weight AI is having its Kubernetes moment
#315Earlier quoted context omitted.
That is already being addressed with GPUs and special PCIe daughter cards that link them together faster than the PCIe on the MOBO could. SLI but better.
Yeah; the benchmarks + setup guide I linked show that it takes a lot of careful setup to get them to work at all, and then you only get 100's of tokens per second per cluster. That's why I pointed out the next generation is coming soon. Also, the AMD docs aren't using quantization (as far as I can tell, I only skimmed), which gives a speedup roughly linear in the compression ratio. Algorithms for that continue to imp…
If you take four Framework Desktop 128 GB Strix Halo motherboards at a previous low and currently-unattainable price, and ignore the cost of storage, power, networking, and cases, and then take an older model that's just about competitive with Opus 4.5, and then you quantise that model down to Q2_K_XL and ignore the performance degradation below Opus 4.5 this will cause, and then you compare the cost of this with the highest tier of Claude subscription that includes much newer models which are far more able than your Opus 4.5 benchmark, then you might have some sort of break-even inside a year.
Yes, there's a new AMD hardware generation coming, but it's going to be as heinously expensive as the current one has become, and it's still got relatively low memory bandwidth. Yes, there are already newer models than Kimi 2.5, but with the limitations of this cluster (e.g. still needing heavily quantised models to be feasible) you'll get incremental improvements at best.
I'm keen to be supportive and bullish about open/home models, but I worry that this is such a stretch, and the two options in your comparison are so incomparable, you'll turn far more people off than you convince.
Re: Open-weight AI is having its Kubernetes moment
#316Earlier quoted context omitted.
GLM 5.2 awq4 via Opencode, it was better than corporate's fave Sonnet 4.8 and a but worse than Opus 4.8. overall quite capable of a lot of what the dev team needed. Not as good as Sonnet 5 but also never runs out of tokens. The flip side is that it takes four H200s to run, and that will only let you cache context for maybe three users. Fingers crossed our Blackwells show up and Kimi 3 really releases weights, because…
What hardware are you expecting to run K3 on?
Re: Open-weight AI is having its Kubernetes moment
#317Earlier quoted context omitted.
It's simple, the US government will put any Chinese open model companies on the entity list which blocks any company which does business with the US from also doing business with the Chinese companies. This creates a chilling effect where even if it may be harder to tell, no US company will be able to provide or use any overt Chinese open model and won't even risk trying to go around as the punishments for trying to…
But that's just the thing with open weights: you're not doing any business with company that made the model. They might publish the weights to a, say, European host, and then you download the model from Europe and and run it on your servers in America, and suddenly it's very hard to tell where the model was originally created.
Re: Open-weight AI is having its Kubernetes moment
#318Earlier quoted context omitted.
> “The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.” That's going to hit first amendment grounds pretty quick, the same way that software in general did. The modern version of the decss flag will be a character that says "I think good weights are {...weights go here...}" They could, however, ban any payment to a chinese entity…
There are already plenty of [illegal numbers]( https://en.wikipedia.org/wiki/Illegal_number ). We have numbers that you can't possess without proper license/authorization, and numbers that you can't yell at a crowded movie theater.
In other words when someone sends you a number that happens to render to CSAM material, its not a coincidence, and any sender pretending it to be coincidence is molesting statistics as well...
Furthermore your "precedent" of illegal numbers is a poor precedent: unlike the tiny fraction of numbers criminalized, all the other numbers remain perfectly legal in stark contrast to big tech proposing to ban all open weight models (!)
Casually dropping in bijections between numbers and images and CSAM images hence CSAM numbers is pure whataboutism that risks ignoring the grave consequences of a ban on open-weight models.
Physics is models. Shall we blanket ban open-source physics?
Hey let's just be silly and casually behave laissez faire when big tech proposes banning open-weight models, because someone has already banned your favorite numbers??!
BTW anyone that follows up with: "hey you only multiplied the appearances of random numbers appearing in communications for a single specific CSAM image, what about the birthday paradox, multiple images are CSAM" don't worry I got you more than covered: 100 bits is not enough to encode a CSAM image worth writing to the police about. An average image easily consumes thousands of bits (or many times more), making the exact number of CSAM images/numbers irrelevant.
Re: Open-weight AI is having its Kubernetes moment
#319I worry about the Chinese models sending data back to China. How do we know that isn't the case?
Re: Open-weight AI is having its Kubernetes moment
#320Everyone just keeps assuming, as if it were the law of gravity, that China will continue in perpetuity to deliver the weights of its 'frontier' models to Hugging Face. Its Mythos moment is a few months away and there is plenty of reporting suggesting their response will be similar, which is anyway obvious. It baffles me that anyone can seriously believe that China is going to put its Mythos successor on Hugging Face…
There is? Xi Jinping himself openly talked very recently about China's commitment to open weights models.