Live data from Hacker News

Outsourcing plus local AI will soon become more economical vs. frontier labs

signalbloom.ai

241–250 of 408 posts

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#241
post #200

I've been saying this for a couple months now since I got decent hardware and started using my local Qwen 3.6 exclusively. I have no doubt the future for individuals and medium-sized companies is local private AI.

Could you share some of your hardware details for Qwen 3.6? And are you using the dense or MoE variant?

Qwen3.6-27B-UD-Q4_K_XL can run at 45t/s with 131k q8 context on an RTX 4090.

That is pretty usable. You could get 65t/s or more with MTP, but only if you drop the context size, which I would advise against.

Results are better with 256k context and a larger quant, however, that's not going to fit on the 4090 you already had lying around for playing cyberpunk 2077.

The MoE models make me rather unhappy. Idk. They feel braindead to me, but YMMV.

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#242
post #169

Earlier quoted context omitted.

https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-...

Does Anthropic really expect to double their income without also doubling their expenses?

I’ve been hearing that anthropic is on the verge of profitability for probably a year straight. Until all the companies agree to stop the training arms race I just don’t see how it’s in the cards

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#243

Earlier quoted context omitted.

There's always wtf, why did we add this feature, but at least in my experience, once a week or so I run into something in this category. Me: "AI, please cleanup/refactor/improve this thing" AI: "Roger that! I deleted the file so now it's perfectly clean" ... insert W.T.F.

Never seen that once.

Same but I have seen it try and change my tests because it decided that it’s code was correct and my tests were incorrect.

I saw it’s thinking tokens said something along the lines of “I have implemented it correctly but the test is failing. I’ll update the tests so the pipeline passes”

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#245

Earlier quoted context omitted.

There are developers outside of your country that are talented, speak your language competently, and willing to work for less pay. There are plenty of reasons to believe that such devs will increase in numbers.

Are they willing to work for less pay than Claude?

Who claimed they would? Cost of labor dwarfs LLM costs.

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#246

I keep seeing this narrative involving Deepseek as an example of OSS LLMs but they are subsidizing a huge amount of tokens at cost and one can easily understand why they are doing it if one is not lazy and think critically. It's still far too costly and not effective to use Local AI that can match what the frontier models can offer, especially when the inference hardware is being heavily restricted due to geopolitica…

Do you have a source for your first claim? My impression is that deepseek designed v4 specifically for cheap inference and they are not loosing money even at 75% lower price.

Do you ? Did you audit deepseek?

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#247

I keep seeing this narrative involving Deepseek as an example of OSS LLMs but they are subsidizing a huge amount of tokens at cost and one can easily understand why they are doing it if one is not lazy and think critically. It's still far too costly and not effective to use Local AI that can match what the frontier models can offer, especially when the inference hardware is being heavily restricted due to geopolitica…

>they are subsidizing a huge amount of tokens at cost This is absolutely false, because other providers serving the Deepseek models on OpenRouter are also able to offer very low prices, and they don't have the money to subsidize anything.

That makes no sense....OpenRouter didn't create Deepseek

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#248
post #157

Earlier quoted context omitted.

Anything over 150 seats means you need to pay at token rates plus the $20/user. My day job is operational (no coding at all) and I'm spending ~$300 a month on a few chats with Claude/Cowork a day over the course of a month.

$300 is my employer's monthly cap on Claude Enterprise. It lasts me at most a week of moderate use. I would much rather get Codex Pro and Claude Pro or Max, which would cost ≤ $200. For $300, one could also add Gemini Ultra to the mix so I could have all three review each other's code, etc. Claude can be very good but enterprise pricing doesn't make sense to me.

That’s a shocking number. I don’t know how much my employer is billed, but based on the numbers reported by Claude code in its optional status bar, I’m often exceeding $300 in a day across sessions, when working on meatier tickets.

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#249
post #79

Earlier quoted context omitted.

I learned today that the Anthropic "Enterprise" plan - the one big companies use because they need governance features and audit logs and all of that jazz - is billed at API token rates (plus $20/seat/month). So large companies are getting billed a lot more than those discount subscription plans.

Anything over 150 seats means you need to pay at token rates plus the $20/user. My day job is operational (no coding at all) and I'm spending ~$300 a month on a few chats with Claude/Cowork a day over the course of a month.

We deployed OpenWebUI with the Claude API the other day for employees. Someone sent ten messages (which appeared to just be reasonable day-to-day work), and we paid $200 for it. There were 44M input tokens, 100k output tokens, no cache hits at all. OpenWebUI reports 3M tokens used, Claude reports 44M, and I have no idea where the rest of the tokens went. This was all on a brand new API key, installed directly to the service, too.

With this kind of opaque billing, how can I reasonably deploy any AI?

Re: Outsourcing plus local AI will soon become more economical vs. frontier labs

#250
post #28

I think this misses the forest for the trees. Working with ChatGPT is eerily similar to working with offshore Indian devs back in my enterprise days. Productive if guided explicitly but if let run wild there's lots of WTF moments. LLMs are likely to replace outsourced devs because your employees that know the context can use LLMs to do what offshore devs did before.

There are developers outside of your country that are talented, speak your language competently, and willing to work for less pay. There are plenty of reasons to believe that such devs will increase in numbers.

yes, every idea's guy that is pumping out SaaS slop as a last ditch effort to avoid a permanent underclass will get priced out of SOTA subscriptions and have to hire cheap offshore developers again. there are a lot of idea's guys. people with no capital and no skills, but ideas.

but for OP's use case, people with some capital and many skills who need additional help, AI is solving a problem in a way that was not solvable before, while improving on coordination abilities and coordination velocity. Offshore developers do not come back into play here.

Post reply on HN