Again, I am talking about inference, not training. Please read.
I indeed failed to understand your point. Isn't that what they already do with Claude Haiku, GPT-5.6-Terra, etc?
That's the goal, but those smaller models don't match the price:performance of leading open-weight models, which is why companies are switching away (as explained in TFA).
The open weight labs figured out some secret sauce that (so far) big name labs are unable to replicate, so instead of competing, they're going on the defensive with claims of distillation attacks.
There are most definitely is a moat - but it works both ways. The railguards in the models create moats keeping customers out. And the cost to build a modern agentic model is in the 10 figure range and growing. This is an expensive arms race that is going to create moats.
> the cost to build a modern agentic model is in the 10 figure range and growing Source? The proliferation of labs building competent models would seem to suggest the opposite.
seems self evident if you read the new. You can Google it yourself, but here's the results from my googling - and this is just for the hardware. Double that to add personnel and corporate infrastructure
"To build or purchase the physical hardware required to store tens of petabytes of data and train a State-of-the-Art (SOTA) frontier AI model, you are looking at a capital expenditure (CapEx) ranging from $320 million to well over $1 billion."
Spent most of today reworking a rack and rig of gpus for all of our internal ai work… our big server is 8 rtx 6000 pro and 3 psu, I definitely feel I made a mistake not upgrading our wall power to 240v but so far we have multiple 30amp 120v and with 3 PSU uninterrupted power we have been very stable. My big upgrade will be moving to epyc motherboard from threadripper so we get gen5x8 with bifurcation instead of what…
What kind of cost are you looking at and how many people could use it? I’d be interested to know what the payback period is like, because Claude code is getting ridiculously expensive.
Bought the rtx 6000 pros one per month starting in January- since they doubled in price I wish I just used a line of credit back in Jan to buy all of them. For me it’s about keeping internal company content internal. Slack channels etc with tools for teams to use. Triage tools for Zendesk etc.
The financials are incredible. If you spend one engineer's salary on hardware, you get a system running local AI that can multiply the efforts of an entire (small) team of engineers. It's a very "you can't afford not to" situation. Even considering the hardware prices today.
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical. Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any ho…
> However the cold reality for both is that there is zero moat to a model anymore. The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability. On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision. On…
You don't need to self host to get the benefits of an open model. There are many hosted providers cheaper than OpenAI or Anthropic who can give you a SLA, ZDR, BAA and all the other three letter acronyms your compliance department needs.
The important part is if they break the contract or raise their prices you can always move to a different provider. You get lower cost and lower risk at the same time which is extremely rare in business. That's just not possible for closed models where your only options are the official branded API or Azure/Bedrock.
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical. Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any ho…
This is exactly true, I get annoyed by Claude one day and switch to something else, and the only thing that's ever keeping me tied towards Claude is the ability to search my old chats easily. But Claude also makes it really hard to do that, so what am I even really paying for? Time to extract all my data, put it into a sqlite with FTS5 and make sure I never rely on the overly-opinionated, low-thinking PMs from these…
Claude by default deletes old chats after a few months. I installed a custom end-of-session hook that throws chat transcripts into a database so that any model can read any other model's chat history. Super easy
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical. Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any ho…
The only moat lives at the Pareto frontier. If you are on the Pareto frontier you are good and can charge money. But the frontier is moving every week so it's super competitive. If you are the quickest innovator, I still believe there is a chance for a working business model for them
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical. Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any ho…
> However the cold reality for both is that there is zero moat to a model anymore. The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability. On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision. On…
Hmm yes but the equipment cost moat is artificial. This scarcity was created by the big AIs by buying up all the future production capacity. That works for a while but it won't last forever.
It's the same with the subscriptions. Local models can't compete because they're simply giving too much value for money. They're effectively subsidised by Big AI. Again something that won't last.
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical. Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any ho…
> However the cold reality for both is that there is zero moat to a model anymore. The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability. On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision. On…
Your annualized energy cost estimates are off by an order of magnitude. 1kwH @ $0.1 (Texas) is $2.40/day if 100% utilized 24x7, California is roughly twice that per my understanding.
Privacy is non-negotiable for corporate. Even without considering costs or country of origin, we've seen from OpenAI that claims of AI safety are worth less than the (virtual) paper they're printed on. All it takes is one incident, and all your company's internal data will start showing up in public users' chats. You can rely on a contract to prevent this, or you can guarantee it by using a locally hosted model you f…
Is this not true for everything then? All cloud services, all Internet providers, everything within M365, OneDrive and Databricks?
Claude is different from S3. AWS doesn’t need to rifle through your files to stay ahead of the competition or to mine them for business ideas because the core business is overvalued and rapidly commoditizing. AI labs, on the other hand, have an incentive to exploit every last drop of data they can lay their ethically challenged hands on. And they’ve already demonstrated that they will do this, even when it involves blatant fiduciary violations (see eg, Anthropic / Figma board member scandal).