One could start with a large model for exploration during development, and then distill it down to a small model that covers the variety of the task and fits on a USB drive. E.g. when I use a model for gardening purposes, I could prune knowledge about other topics.
Small language models are the future of agentic AI
21–30 of 47 posts
Re: Small language models are the future of agentic AI
#22A few weeks ago, I processed a product refund with Amazon via agent. It was simple, straightforward, and surprisingly obvious that it was backed by a language model based on how it responded to my frustration about it asking tons of questions. But in the end, it processed my refund without ever connecting me with a human being. I don't know whether Amazon relies on LLMs or SLMs for this and for similar interactions,…
You can (used to?) get a refund on Amazon with normal CRUD app flow. Putting an SLM and a conversational interface over it is a backwards step.
We're going to be so messed up in a decade or so when only 10-20-30% of the population is employable in decent jobs.
People keep harping on about people moving on with their lives, but people don't. Many industrial heartlands in the developed world are wastelands compared to what they were: Walloonia in Belgium, Scotland in the UK, the Rust Belt in the US.
People don't really move on, they suffer, sometimes for generations.
Re: Small language models are the future of agentic AI
#23This is not a minor oversight - it's arguably, in my experience, the most prohibitive technical barrier to this vision. Consider the actual context requirements of modern agentic systems:
- Claude 4 Sonnet's system prompt alone is reportedly roughly 25k tokens for the behavioral instructions and instructions for tool use
- A typical coding agent needs: system instructions, tool definitions, current file context, broader context of the project it's working in. Additionally, you might also want to pull in documentation for any frameworks or API specs.
- You're already at 5-10k tokens of "meta" content before any actual work begins
Most SLM that can run on consumer hardware are capped at 32k or 128k contexts architecturally, but depending on what you consider a "common consumer electronic device" you'll never be able to make use of that window if you want inference at reasonable inference speeds. A 7b or 8b Model like DeepSeek-R1-Distill or Salesforce xLAM-2-8b would take 8GB of VRAM at Q4_K_M Quant with Q8_0 K/V cache at 128k context. IMO, that's not just simple consumer hardware in the sense of the broad computing market, it's enthusiast gaming hardware. Not to mention that performance degrades significantly before hitting those limits.The "context rot" phenomenon is real: as the ratio of instructional/tool content to actual tasks content increases, models become increasingly confused, hallucinate non-existent tools or forget earlier context. If you have worked with these smaller models, you'll have experienced this firsthand - and big models like o3 or Claude 3.7/4 are not above that either.
Beyond context limitations, the paper's economic efficiency claims simply fall apart under system-level analysis. The authors present simplistic FLOP comparisons while ignoring critical inefficiencies:
- Retry tax: An LLM completing a complex task with 90% success rate might very well become 3 or 4 attempts at task completion for an SLM, each with full orchestration overhead
- Task decomposition overhead: Splitting a task that an LLM might be able to complete in one call into five SLM sub-tasks means 5x context setup, inter-task communication costs, and multiplicative error rates
- Infrastructure efficiency: Modern datacenters achieve PUE ratios near 1.1 with liquid cooling and >90% GPU utilization through batching. Consumer hardware? Gaming GPUS at 5-10% utilization, residential HVAC never designed for sustained compute, and 80-85% power conversion efficiency per device.
When you account for failed attempts, orchestration overhead and infrastructure efficiency, many "economical" SLM deployments likely consume more total energy than centralized LLM inference. It's telling that NVIDIA Research, with deep access to both datacenter and consumer GPU performance data, provides no actual system-level efficiency analysis.For a paper positioning itself as a comprehensive analysis of SLM viability in agentic systems, sidestepping both context limitations and true system economics while making sweeping efficiency claims feels intellectually dishonest. Though, perhaps I shouldn't be surprised that NVIDIA Research concludes that running language models on both server and consumer hardware represents the optimal path forward.
Re: Small language models are the future of agentic AI
#24A few weeks ago, I processed a product refund with Amazon via agent. It was simple, straightforward, and surprisingly obvious that it was backed by a language model based on how it responded to my frustration about it asking tons of questions. But in the end, it processed my refund without ever connecting me with a human being. I don't know whether Amazon relies on LLMs or SLMs for this and for similar interactions,…
I just had my first experience with a customer service LLM. I needed to get my account details changed, and for that I needed to use the customer support chat. The LLM told me what sort of information they need, and what is the process, after which I followed through the whole thing. After I went through the whole thing it reassured me everything is in order, and my request is being processed. For two weeks, nothing…
Re: Small language models are the future of agentic AI
#25Earlier quoted context omitted.
You can (used to?) get a refund on Amazon with normal CRUD app flow. Putting an SLM and a conversational interface over it is a backwards step.
From our perspective as users. From the company's perspective? Net positive, they don't need to hire people. We're going to be so messed up in a decade or so when only 10-20-30% of the population is employable in decent jobs. People keep harping on about people moving on with their lives, but people don't. Many industrial heartlands in the developed world are wastelands compared to what they were: Walloonia in Belgiu…
The LLM, here, is the opposite; additional human labor to build the integrations, additional capital for chips, heavy cost of inference, an additional skeuomorphic UI (it self identifies as a chat/texting situation) and your wasted time. I would almost call it "make work".
Re: Small language models are the future of agentic AI
#26A few weeks ago, I processed a product refund with Amazon via agent. It was simple, straightforward, and surprisingly obvious that it was backed by a language model based on how it responded to my frustration about it asking tons of questions. But in the end, it processed my refund without ever connecting me with a human being. I don't know whether Amazon relies on LLMs or SLMs for this and for similar interactions,…
I just had my first experience with a customer service LLM. I needed to get my account details changed, and for that I needed to use the customer support chat. The LLM told me what sort of information they need, and what is the process, after which I followed through the whole thing. After I went through the whole thing it reassured me everything is in order, and my request is being processed. For two weeks, nothing…
Re: Small language models are the future of agentic AI
#27How is SLM the future of AI while we are not even sure about if LMs are the future of AI?
Re: Small language models are the future of agentic AI
#28A few weeks ago, I processed a product refund with Amazon via agent. It was simple, straightforward, and surprisingly obvious that it was backed by a language model based on how it responded to my frustration about it asking tons of questions. But in the end, it processed my refund without ever connecting me with a human being. I don't know whether Amazon relies on LLMs or SLMs for this and for similar interactions,…
You can (used to?) get a refund on Amazon with normal CRUD app flow. Putting an SLM and a conversational interface over it is a backwards step.
The product maker cecotec, from whom I hope never to buy a product from, uses the following repairing process: you have to create an account to submit your data and then when you try to login into the caretaker page the system announces that the account you created does not exist (I have tried several times, several days, both with mine and my wife email and personal data). Furthermore there is no way someone should take the telephone on cecotec.
Another step that failed is to ask the Spanish postal service to give me some option to try to send again the product from its current location, stopped expecting sender order, to Amazon Storage avoiding the product come back step: they informed me that they can not respond to my email because the privacy law forbid it, perhaps this is because the product was send with my wife name and address. I don't suppose the air conditioned system to have some personal information attached to it.
Buddishm says that you can learn from any experience and maybe the karma of this product is to never be repaired and the fate has decided to condemn the old lone woman to suffer the hot wave, perhaps to expire some past life bad karma.
Fortunately all that comes goes by, and in this case I am happy to be able affording to buy a new machine that produces cold air, so, kind reader, be quiet there is no real problem. Furthermore, this experience can expand my empathy: it could be that for some people life don't work as it should, for them this anecdote is the normal course of actions where one problem calls another. For those a stream of problems is the only repl. To those I, most sincerely, wish peace of mind and hope their fortune reverse.
In this anecdote or episode human intelligence is not producing correct results and that could create a hope: that those small LLM models enhanced with intelligent agents could provide better support.
Today, in my current mental state I envision that the contrary could occur: those system could convert the bad things that happens once into the bad things that happens every day. So be careful with what you wish.
My hope it is that the greatest agents, ourselves, get together to solve whatever problem we have to cope with. But don't fool yourself that require real human deep intelligence and human hard work.
Re: Small language models are the future of agentic AI
#29Anyway really love the idea but many years of experience with decentralization / security / privacy projects makes me think it probably won't happen. Their description of how to incorporate SLMs at the end gives the game away: it's a description of a large, complex project that requires fairly good data science skills e.g. they just casually suggest you autoclean the data of PII then run unsupervised clustering over the results to prepare data for model fine tuning using QLoRA, and then set up an automated pipeline to do this continuously. Sure. We'll get right on that.
The history of computing is pretty simple: given a choice between spending more on hardware or more on developers, we always prefer more hardware. For NVIDIA this is a good thing modulo the fact that nobody buys their hardware directly because it's too overpowered. But that's the way they've chosen to segment the market. Given a choice between using a sledgehammer to crack a nut, or to make a custom hammer for nut cracking, we're gonna spend the money on outsourced LLMs every time.
Naturally, perhaps in future LLMs will create these SLM factories for us! If you assume software is mostly written by LLMs in future then past experience about expressed preference of software teams might not apply. But we're not there yet.
Re: Small language models are the future of agentic AI
#30Wonder what I'm missing here. A smaller number of repetitive tasks - that's basically just simple coding + some RPA sprinkled on top, no? Once you've settled down on a few well-known paths of action, wouldn't you want to freeze those paths and make it 100% predictable, for the most part?
These sorts of “heuristics are surprising incapable while the semantic flexibility of language models are powerful” are surprisingly large. Even flexible validation mechanisms that take human entered semi fixed form inputs and reformat them to the expected input in a really reliable way and is much less frustrating to end users. Essentially any situation where abductive logic or natural language comes into the picture a small model does really well.
Both of those things - abductive logic and natural language - were largely unavailable as tools until recently. This pretty nearly rounds out the complete toolkit for making really robust and powerful (and usable) systems. You sacrifice a perceived determinism by admitting abductive logic and non determinism, but in my experience this warm blanket of the mathematically inclined wasn’t particularly robust in reality and systems often deterministically failed in complex and difficult if not impossible ways to avoid that a little bit of abductive reasoning could make it remarkably simpler and more robust.