Earlier quoted context omitted.
I run it on my 4 year old MBP and get 10 tok/s. With the RAM shortage buying anything new today is a nightmare but anyone with a reasonably modern Mac could run it at q6 probably. It is mostly a toy as 4o models weren’t really suitable for real work IMO but at least it won’t ever give me a refusal.
At 10toks, are you using it interactively or do you submit a prompt and come back to it later? I always thought it would make sense to just do conversations over email, asynchronously, the model can take all the time it needs and get back to me when it has an answer.
AI subscriptions are a ticking time bomb for enterprise
261–270 of 426 posts
Re: AI subscriptions are a ticking time bomb for enterprise
#262Earlier quoted context omitted.
"It costs OpenAI less money to serve GPT-5.5 than GPT-4." does it though? do you have the numbers? Or you just making stuff up?
GPT-4 (original API): Input: $30 / 1M tokens Output: $60 / 1M tokens GPT-5.5: Input: $5 / 1M tokens Output: $30 / 1M tokens Costs have been reducing by over 5x year over year. Inference cost concern is mostly performative. https://simianwords.bearblog.dev/conclusive-proofs-that-llm-... Edit: can't reply but companies aren't selling inference at loss. In the blog post I point to third party hosting of open models like…
We will only know the actually situation once Anthropic goes public and we can look at their books.
Re: AI subscriptions are a ticking time bomb for enterprise
#263Earlier quoted context omitted.
> Local modals are 6 months to 18 months behind frontier. I wish this was true but it is not. And I am working on open source models so if anything, I would have a bias towards agreeing with you. Frontier closed models (GPT/Claude) are gaining distance to everybody else. Even Google, once the king. Your claim is a meme coming from benchmark results and sadly a lot of models are benchmaxxed. Llama 4, and most notably…
The Chinese models should stay close on a lag. They’re doing a ton of distillation that, realistically, I’m not sure the American frontiers can stop.
[0] US AI firms team up in bid to counter Chinese 'distillation' (Apr 7) https://finance.yahoo.com/sectors/technology/articles/us-ai-...
Re: AI subscriptions are a ticking time bomb for enterprise
#264Earlier quoted context omitted.
"It costs OpenAI less money to serve GPT-5.5 than GPT-4." does it though? do you have the numbers? Or you just making stuff up?
GPT-4 (original API): Input: $30 / 1M tokens Output: $60 / 1M tokens GPT-5.5: Input: $5 / 1M tokens Output: $30 / 1M tokens Costs have been reducing by over 5x year over year. Inference cost concern is mostly performative. https://simianwords.bearblog.dev/conclusive-proofs-that-llm-... Edit: can't reply but companies aren't selling inference at loss. In the blog post I point to third party hosting of open models like…
That blog post is not very compelling either. Without knowing details of the architecture, comparing the various frontier models to open models doesn’t make sense.
Re: AI subscriptions are a ticking time bomb for enterprise
#265[flagged]
Re: AI subscriptions are a ticking time bomb for enterprise
#266Earlier quoted context omitted.
"It costs OpenAI less money to serve GPT-5.5 than GPT-4." does it though? do you have the numbers? Or you just making stuff up?
GPT-4 (original API): Input: $30 / 1M tokens Output: $60 / 1M tokens GPT-5.5: Input: $5 / 1M tokens Output: $30 / 1M tokens Costs have been reducing by over 5x year over year. Inference cost concern is mostly performative. https://simianwords.bearblog.dev/conclusive-proofs-that-llm-... Edit: can't reply but companies aren't selling inference at loss. In the blog post I point to third party hosting of open models like…
Pricing has no correlation with profit. It can be artificially lowered to kill competition, and artificially inflated to maximize profit.
Re: AI subscriptions are a ticking time bomb for enterprise
#267Enterprise customers aren't running 20 bucks a month for claude pro subscriptions. My company provides developers about 1k worth of usage limits a month and best I can tell they get maybe a 30% savings off of API cost tops. That's not an insane subsidy. Many other jobs titles are only allowed 50 a month and those folks are constantly running out. Github Copilot has been doing this with business and enterprise seats,…
I was actually quite worried, because I've been using GHCP for large chunks of work, but the billing estimator they released shows I was only at about $150-200 a month in API priced tokens. Sure, that's a subsidy for my $20 subscription, but not insane.
Heavy use of agentic coding tools, in a responsible manner, probably lands somewhere around that $200/m mark at API pricing. Assuming that makes the provider money, I don't see that being hard to swallow for businesses employing developers in Western countries, given the hours it can save.
The real risk here is to personal project vibe coders. Building a huge app by abusing subsidized plans is ending.
Re: AI subscriptions are a ticking time bomb for enterprise
#268Earlier quoted context omitted.
Something I have noticed is that the people who are using it to write everything are the same people who had a poor level of English writing a year or two ago. It's just "intellectual" botox.
> It's just "intellectual" botox. Could be just ESL, it's hard to close the proficient to native gap.
Re: AI subscriptions are a ticking time bomb for enterprise
#269There will be a repricing for sure as any ends of subsidies does but the world will not end
Re: AI subscriptions are a ticking time bomb for enterprise
#270Every AI subscription is a ticking time bomb for the frontier provider; within a few years we will be running local models as good as today’s frontier models with almost no cost burden. The floor will fall out of the enterprise market for all the frontier companies.
> within a few years we will be running local models as good as today’s frontier models with almost no cost burden Based on what? The RAM requirements alone are extraordinary. No, running large models on shared, dedicated hosted hardware at full utilization is going to be vastly more cost-efficient for the foreseeable future.
I take it you haven’t actually run any of the current gen local models?
They all fit on fairly accessibility hardware, and their performance is at least on par with what I was paying for last year.
I have one of my agents running entirely from a local model running on a MBP and it has repeatedly shown it’s capable of non-trivial tasks.
Playing around with another, uncensored, local model on my 4090 desktop has me finally thinking about canceling my personal Anthropic subscription. Fully private, uncensored chat is a game changer.
For work it’s still all private models but largely because, at this stage, it’s worth paying a premium just to be sure you’re using the best and it saves the time of managing out own physical servers. But if we got news tomorrow that Anthropic and OpenAI were shutting down, a reasonable setup could be figured out pretty quickly.