Earlier quoted context omitted.
Incentives matter… If prices keep going up, watch for companies to exit frontier models and go to local llama.cpp instances for 6-month-ago SOTA, with the flex of being housed within the office - no more privacy leakage, no more price gouging. To be honest, I’m not sure why a Y-Combinator backed company hasn’t come out yet flooding the market with highly capable OPAI (pronounced “Oh-pah” as in what Greeks shout as th…
> If prices keep going up, watch for companies to exit frontier models and go to local llama.cpp instances for 6-month-ago SOTA, with the flex of being housed within the office - no more privacy leakage, no more price gouging. That or just hiring people to do the work! I hear rumours that this is already starting to happen in some places (perhaps those that were a little overzealous with AI-hype driven layoffs).
Outsourcing plus local AI will soon become more economical vs. frontier labs
251–260 of 408 posts
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#252Earlier quoted context omitted.
Right.. and their definition of a "best product" is theirs, not mine and probably not yours.
…the best model for agentic coding, is the top goal right now.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#25327B is already really good at coding-specific tasks. Fundamentally, there is little innovation on the core architecture: LLMs are all designed essentially the same, with minor differences in how they are trained. They are all feed-forward multi-headed attention models; it doesn't matter if it's a 4B model or a 1T model, that's just scale.
Further, the frontier models cannot afford to innovate: they have to scale as quickly as possible to "beat out" their competition. The frontier models fundamentally will not create the next "attention is all you need" monumental jump in AI.
Frontier companies are stuck on scale with zero capacity to innovate. You cannot point capitalism at "basic science research" and expect any ROI. This is a known reality. Innovation is much more indirect and a "random walk" style of knowledge acquisition.
Finally, these LLMs are quite literally designed with a human-in-the-loop, and we do not give ourselves enough credit for how well we ourselves tool-call. We are doing a lot of heavy lifting to make these models useful and you cannot simply remove us from the equation without also removing ourselves from the training pipieline.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#254Earlier quoted context omitted.
Depends on what you mean by "economically feasible". Even very cheap mini-PCs and laptops can run any of the models run by cloud providers, albeit at a much lower speed (i.e. with the weights stored on SSDs). Whether such a low speed is useful, depends on the application. For something like a coding assistant or bug scanning, an instant response is desirable, but certainly not necessary.
The SSD would wear out in days while the laptop generates two responses a day. This is like saying you could power your home with AA batteries, yes technically you could but in practice entirely infeasible.
For model training, the requirements are very different, and the training of a big LLM cannot be done with home equipment. On the other hand, inference can be done on almost any PC, even for LLMs with thousands of billions of parameters, just very slowly.
The only problem is that the inference becomes limited by the SSD reading throughput. Most of the cheap new personal computers available today can read simultaneously only 2 SSDs (if there are more they share a reading path), which are typically 1 PCIe 5.0 SSD and 1 PCIe 4.0 SSD. This has an upper throughput limit of 24 Gbyte/s, with 15 to 20 GB/s achievable in practice.
Then the speed in token/s is limited by the amount of weights that must be read per inference cycle. The ratio between output tokens and the amount of weights that must be read can be improved by various methods, like batching multiple tasks or using speculative decoding.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#255Earlier quoted context omitted.
"roughly" is doing a lot of heavy lifting there
The difference between datacenter hardware and cheap personal hardware is not in what can be run and what cannot be run. Anything can also be run on a cheap computer. The difference is in speed. A cheap computer may run a big model up to a few orders of magnitude slower than datacenter hardware, depending on whether the LLM is small enough to fit in GPU memory, or it is small enough to fit in CPU memory or it is so b…
You do realize that a model like Opus is (estimated to be) around 5T parameters, and uses around 5TB of GPU memory?
These kind of things are just impossible to run locally.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#256Earlier quoted context omitted.
I used to be on 5.4 high for most of my work. I have switched completely to 5.5 medium now. I would highly recommend trying it out - 5.5 is significantly more token efficient than 5.4 - the same task takes often a third of the tokens - because of this, is it also much faster to do the task - you get high "intelligence" per token even after accounting for token efficiency - 5.5 medium is just under 5.4 pro levels of i…
This is embarrassing but I find 5.4-mini on Low covers a substantial part of my and my colleagues work. Back when it became expensive I learned to live with it and I find my "AI skills" (mainly communication) have a substantial impact on the efficiency of the model. Not saying my work is difficult, it's not, but I find there is quite a bit of wiggle room. Smaller models can still perform useful work, but you have to…
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#257Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#258The weaker a developer is the higher capability AI requires. The entire premise of this article does not work because it confuses weak developers with weaker ai being better than strong developers with near atonomous ai. The weak developers with frontier ai already produce products that are worse than a capable developer paired with a weak (2 year old) AI.
To clarify: Strong developers 2 years ago could already leverage AI to produce high quality products whereas with latest and greatest AI weaker developers stills struggle strong developers can now delegate more of the work to the stronger AI increasing productivity further.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#259Earlier quoted context omitted.
The difference between datacenter hardware and cheap personal hardware is not in what can be run and what cannot be run. Anything can also be run on a cheap computer. The difference is in speed. A cheap computer may run a big model up to a few orders of magnitude slower than datacenter hardware, depending on whether the LLM is small enough to fit in GPU memory, or it is small enough to fit in CPU memory or it is so b…
> The difference between datacenter hardware and cheap personal hardware is not in what can be run and what cannot be run. You do realize that a model like Opus is (estimated to be) around 5T parameters, and uses around 5TB of GPU memory? These kind of things are just impossible to run locally.
Like I have said, the problem is not that they cannot be run, but that they may run more slowly than it is acceptable for a given application. Depending on the model, the speeds reported for inference with weights stored on SSDs vary from one token every few seconds to at most a few tokens per second.
Computers could solve relatively huge problems even in the early days of vacuum tube computers, when the main memories were measured in kilobytes, because at that time it was not expected that the data needed for problem solving must fit inside the main memory or even in the next tier of memory, with magnetic drums or magnetic disks, but the really big problems were solved by a great number of passes over data stored on magnetic tapes.
An LLM whose inference could not be run on a small mini-PC would have to be one hundred times bigger than the biggest existing SOTA LLMs.
Any LLM that exists today can be run on almost any PC, just extremely slowly in comparison with datacenter hardware.
Re: Outsourcing plus local AI will soon become more economical vs. frontier labs
#260Earlier quoted context omitted.
$300 is my employer's monthly cap on Claude Enterprise. It lasts me at most a week of moderate use. I would much rather get Codex Pro and Claude Pro or Max, which would cost ≤ $200. For $300, one could also add Gemini Ultra to the mix so I could have all three review each other's code, etc. Claude can be very good but enterprise pricing doesn't make sense to me.
The $200 plan you're talking about is subsidized by Anthropic. They cannot afford to keep offering that to everyone indefinitely. Absolute best case scenario for current users is that they can continue to subsidize it as way to sell enterprise plans, but there's no way that they can keep offering it to everyone at those prices.
Common talking point. There's enough evidence for the counter argument that this is essentially misinformation. I have no idea why it's so often repeated with confidence.