Earlier quoted context omitted.
I'm sure there's some good ones but most are bad. Not because all Indian devs are bad (this is of course absolutely not true), but most of the good ones are either no longer in India or working in India but for something more prestigious and interesting than an offshoring shop
Which can be said of any country. Most of the good devs in US are working for the best US comapnies and not for small companies with less budgets.
Outsourcing plus local AI will soon become more economical vs. frontier labs
341–350 of 408 posts
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#342Earlier quoted context omitted.
Sure, I have a 64G MBP with an M1 Ultra. The best model for me by far has been the 35B A3B, in particular the 8Q_KL unslouth variant. The dense model works but it's much slower, and I don't really see a difference in quality with a good harness.
What do you use as a harness?
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#343Earlier quoted context omitted.
Do people commonly argue Waymo isn't subsidizing rates? Also, we do have some evidence for my position: - We know that the consumer Claude plans provide _way_ more tokens than you could get if you were paying API prices. This is a huge part of why Anthropic's limits on other harnesses for subscription customers is such a big deal. So either their profit margin on API tokens is absurdly high, most consumer subscribers…
> I think the case that their consumer subscriptions lose them money on net is pretty strong, even though their enterprise subscriptions (and API pricing) does make them a profit. To be clear I'm not arguing against this position, just questioning the confidence with which people claim that the current consumer subs are not a sustainable offering and a merely temporary.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#344Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#345Earlier quoted context omitted.
> Lastly, there is a massive difference in capabilities, determinism, and error handling between 5T SOTA models like Opus What's your source for Opus being a 5T model? > and tiny distillations from DeepSeek that perform well only in benchmarks. I don't think you know what you're talking about. Local models aren't “distillations from Deepseek”. And they don't perform well “only in benchmarks”, Qwen 3.6 is a very decen…
> What's your source for Opus being a 5T model? Elon Musk tweeted that Grok is 0.5T or 1/10th the size of Opus. https://xcancel.com/elonmusk/status/2042123561666855235#m While this source's reliability is certainly debatable, the size matches the results of this paper, in which researchers estimated the parameter count from model knowledge. https://01.me/research/ikp/
It's certainly a better sell that Grok sucks because it's small and Opus is impressive because it's large, than the alternative that Grok is also large and sucks which points to xAI incompetence and mismanagement.
Particularly when you're trying to IPO a rocket company based on rosy forecasted valuations of Grok dominating the market.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#346Earlier quoted context omitted.
Most software development teams are pushing back on the deluge of bad changes from AI tools and are moving slower again to regain trust and stability. It is likely that future software development will not actually be higher velocity.
Bad changes will be eliminated because better people are using the AI tools. They will reduce cost as well as slop.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#347Earlier quoted context omitted.
> What's your source for Opus being a 5T model? Elon Musk tweeted that Grok is 0.5T or 1/10th the size of Opus. https://xcancel.com/elonmusk/status/2042123561666855235#m While this source's reliability is certainly debatable, the size matches the results of this paper, in which researchers estimated the parameter count from model knowledge. https://01.me/research/ikp/
Elon Musk has absolutely no credibility anymore. I'm more likely to believe the opposite of what he claims to be true.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#348Earlier quoted context omitted.
Sell more. The hope is that there is a huge addressable market that includes huge per-worker demand in almost all white collar work and lots of inference in people's private lives If that doesn't work, then yes, then prices will have to go up
Both anecdotally for myself and from what I'm reading in the news, it seems just as likely that AI usage has already largely peaked. There was a lot of hype and exploration of capabilities, but models aren't evolving fast enough to keep that going, so I'm settling down into a familiarity with what an LLM can and can't do that means I am using them less overall that I was 6 months ago when I was throwing everything un…
I think this is minimally likely. While as individuals on the bleeding edge, we're perhaps using these tools less and less, and our echo chamber reinforces that, the penetration of AI into the normal corporate workplace is still very low - emails rewritten with ChatGPT, meeting notes summaries generated by default, etc. There are a million use cases for LLMs which are not yet built out. The tokenmaxxers will begin using AI less, but the penetration into the mass market will continue at a huge velocity.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#349If you value high quality work and pride in what you do, outsourced workers (who most often don't pay careful attention to their work, hence the cost) are not the solution. However, if you're just trying to get something done and don't care about it getting done right, what better way to do it than spending the least amount of money possible
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#350Earlier quoted context omitted.
When people say that you "can't do" something what they actually mean is that it's completely impractical (if not impossible).
Whether something is "impractical" depends on your expectations. High-latency unattended inference is definitely viable, even though it doesn't align much with what's being run in hyperscale datacenters.
I think 1 token/second is optimistic here - and even then it's over 11 days per million tokens.