Live data from Hacker News

Open-weight AI is having its Kubernetes moment

tobi.knaup.me

141–150 of 346 posts

Re: Open-weight AI is having its Kubernetes moment

#141
post #2

One of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything. So what open weight models do is at least provide a baseli…

What is strange about GPT-4 being expensive in 2023? Supply and demand. Which other model choices did we have? Prices are related to supply and demand. We see it play out with the introduction of capable open weight models or even other closed cloud models.

Not really though right?

As of early June 2026, Opus 4.8 in fast mode cost $50/M output tokens and Opus 4.6 & 4.7 cost $150/M output tokens in fast mode

How can supply and demand explain the price drop? Was it cheaper to serve Opus 4.8? Is the demand for the newer Opus lower than for the older Opus? These are just fixed prices that seem picked out of thin air

Re: Open-weight AI is having its Kubernetes moment

#142
post #16

Is anyone using open weight models for agentic coding? What is your stack (harness, model) and how much do you pay per month? How would you compare your experience to a typical subsidized plan like Claude Code + Pro plan? I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?

I'm using Qwen 3.6 27B on a macbook with Pi. It's alright, it runs fairly quick (40 tps for quality version, 80 for the fast). It doesn't tend to one shot things but I'm generally comfortable fixing the bugs myself afterwards or prodding it a little bit. I find the harness matters a lot. "Continue" (the vscode extension) worked horribly, OpenCode was ok but its vibecoded internals make me view it as a security nightmare so I'm hesitant to run it, so I've settled on Pi for now.

Claude and ChatGPT are good deals right now, with the subsidies. They produce things faster and better. I guess not cheaper, in that inferrence on my macbook is basically free, although the macbook itself definitely wasn't. My focus on running local is around three principles:

1. I don't want to support surveilance capitalism by giving these companies my data anymore, when I can avoid it. And LLM companies want to vacuum up every detail of your life.

2. I don't find these companies to be remotely trustworthy, and I find them hostile to a healthy society, so I want to avoid giving them money going forward

3. I think they're going to start charging a lot more

Re: Open-weight AI is having its Kubernetes moment

#143

Earlier quoted context omitted.

Rex 3090 isn’t free. Even if you already owned it, it wasn’t free. That’s years of a $20 subscription

It's an asset, though, and bizarrely it's one that's been appreciating the past 5 years

I took the same attitude. The hardware isn’t getting cheaper, it’s getting more expensive.

As I see it, an investment in AI hardware is an investment in my own future.

IE, I drive my car a couple of days a week, and it’s perfectly normal to spend $500 a month on an asset like that. When you factor in the SPACE it takes up, that’s the REAL cost of owning a car: the real estate you have to buy for your car to occupy.

Once that’s factored in, the “true” cost of having a car can easily be $2000 a month, even for a crummy car. The space that the car occupies is expensive.

Yet people balk at spending even $2000 on a GPU.

Makes no sense to me. I choose to invest in the future.

Re: Open-weight AI is having its Kubernetes moment

#144
post #135

Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself. It is good it exists though to put pressure against the labs. Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.

Pre bubble prices (~= “we stop building data centers with subsidized credit / circular loans / hidden debt”), a 128GB halo strix ran for $1400, and 200-ish watts. Four of those in a cluster will run a 1T parameter frontier model: https://www.amd.com/en/developer/resources/technical-article... At 7 months of claude code subscription per node, the cluster pays for itself in 28 months. On a 5 year (60 month) depreciatio…

[deleted]

Re: Open-weight AI is having its Kubernetes moment

#145

Earlier quoted context omitted.

> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. The price is what the market is willing to bear for the available compute capacity and competitive landscape. You can only discover that price after trying different price points and seeing what happens. Everyone is trying different pricing schemes and discounts as th…

Most unmature markets aren't subsidized to the point that LLM market is, most market have some level of baseline profitablity, this market doesn't, that's because most market subsidized the marketing or the capex but this market doesn't hold the opex, the capex not the amount of marketing let alone all of this together

[dead]

Re: Open-weight AI is having its Kubernetes moment

#146
> American labs need to release frontier-grade open-weight models under licenses that startups can actually build on.

To be fair, OpenAI has released a couple of (then very good) OSS models. I run the 20B version at home and it is excellent for reviewing text and common tasks like drafting bash scripts. There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too. I wish OpenAI updated these models more frequently though.

Re: Open-weight AI is having its Kubernetes moment

#148

Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself. It is good it exists though to put pressure against the labs. Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.

Model-on-Chip is coming. GPU are for general computing but have a huge bottle neck for doing model inference. Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.

This will probably work well with SotA planning and local chip implementation. I see them being like cars. Cost a few 10k on credit, buy a new one when the old one goes bad or marketing convinces you to upgrade.

Re: Open-weight AI is having its Kubernetes moment

#149
post #126

Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive at…

You need to ask what happened in Tienanmen square

Re: Open-weight AI is having its Kubernetes moment

#150
post #82

FTFA: American labs need to release frontier-grade open-weight models under licenses that startups can actually build on. oh now i see, the Chinese government is funding the training and release of their best models to pressure OpenAI, Anthropic, and others to do the same for competition's sake. I don't buy it, this seems more like a way to get SOTA models RL'd to comply with Chinese government approved information d…

What tools do we have to countermeasure the state sponsored bias in the Chinese models? Doesn’t seem like a smart plan if individuals can just compensate for the bias.

Also, the choice right now is between an open weight Chinese model that is hypothetically censored to block / sabotage routine engineering flows vs a closed weight service that is definitely censored to block / sabotage those things.

First anthropic guardrails blocked totally normal stuff on fable and knocked you down to opus. At this point, they kick you off fable, then opus, then sonnet. Claude then automatically builds up memories of techniques to bypass the guardrails in my long running sessions (the coordinator agent notices the subordinates got shot in the head and their sessions were pulled from context, so it parses out the lost context from ~/.claude json files, then reformulates parts of the task and uses partial results until the guardrail doesn’t trip.

If I were paying for the API, this dance would cost $50-100 a pop, but I’m not, so whatever (for now).

Post reply on HN