Live data from Hacker News

Everyone is building LLM routers, we deprecated ours

manifest.build

31–40 of 94 posts

Re: Everyone is building LLM routers, we deprecated ours

#31
I agree with one distinction - coding agent workflows can use defined subagent roles that are pinned to specific models and I have found this very effective. The orchestrator is building all the context to make these assignments - it’s not a dumb router. Using Minimax M3 for exploration and librarian tasks for example is fast and cheap - my $10 plan lasts all month and saves a lot of tokens for my main coding plan.

Re: Everyone is building LLM routers, we deprecated ours

#32
post #23

Earlier quoted context omitted.

You also have problems with short prompts whose results depend highly on the understanding of nuance. The router is going to have to mostly solve the prompt to decide which model to send it to. What you actually want is a model that can conclude either "I know the answer to this with confidence" and answer, or "I think I don't know the answer to this, I should ask another model and I know which one". But I don't thin…

> Their internal confidence can be measured and returned It can? I was under the impression that confidence was either self-reported by the LLM or assessed by having another model interpret the output response. If there's a confidence score at the level of the actual model math, that's news to me.

You certainly can't ask an LLM for its confidence level because it will generate that answer as it sees fit.

My understanding was that models can be post-trained to assess their own confidence on short prompts within a level of accuracy (there was an article about this a few weeks ago that I can't find), but can't use it to reason, and that research has shown that internally they effectively have a measurable sense of truth but can't surface it:

https://arxiv.org/abs/2410.02707

I mangled what I was trying to say but the point I guess is, it is effectively in there, and researchers can see it and therefore score it, but it is not something the LLM can use.

The problem seems to me (layman's understanding at best) that the LLM is inherently confident within the words of its answer, because of what the model is trained to do and how it is trained. So you can never get an accurate "I don't know this" from a model while it is answering; it is bullshitting.

(One of the things that is most interesting to me at the moment is asking a small local LLM what it knows about a topic and then testing it. It can have no idea of what it wasn't trained on, but any question about what it knows is seemingly going to trigger it to work through all the summary word associations it did find in training, so it can sort of summarise its knowledge that way, with some likelihood of success, without any sense of introspection)

Re: Everyone is building LLM routers, we deprecated ours

#33
post #29

Earlier quoted context omitted.

> Their internal confidence can be measured and returned It can? I was under the impression that confidence was either self-reported by the LLM or assessed by having another model interpret the output response. If there's a confidence score at the level of the actual model math, that's news to me.

At least one model can today. https://github.com/cactus-compute/cactus-hybrid

Ahh you found the link I couldn't find with a search.

Yes — they claim that it is likely to be able to assess confidence in whether the statement it just made is correct. But it isn't going to be capable of that while it is answering.

Re: Everyone is building LLM routers, we deprecated ours

#34
Insider take: routing will not be a (successful, durable) thing, at least not externally to model providers.

The labs are incentivized to solve this problem themselves, since they’re competing on a 2D cost-intelligence frontier. If they can reduce cost without harming intelligence they will do that and pass on (some of) the cost reduction to the user. There are nicer solutions available to them because they can cut into lower levels of abstraction.

E.g. you should consider speculative decoding to be one (very conservative) form of routing and note that you can’t implement that for the labs from the outside.

Re: Everyone is building LLM routers, we deprecated ours

#36

> Just as a painter knows exactly what brush they need to use, and the craftsman carefully chooses their tools, engineers should understand trade-offs and subtleties of the different models. I'm really skeptical of this idea. Pragmatically: who has time to understand the nuances of these models when there's like a new one every week? Also without any view into the training, figuring out what each model is potentially…

I'm not surprised that people who use LLMs all day for everything get a feel for the different ones.

Re: Everyone is building LLM routers, we deprecated ours

#37
post #4

Ironically, my confidence that a human had at least an active part in writing/editing this article went up because of this train wreck of a sentence: > "A cache-aware model router will take that into account by adding stickiness to the initially chosen model and keeps querying it."

what’s wrong with the sentence? reads fine to me

Weird tense combination. I'm not sure if it's officially wrong but it feels weird to change tense mid-sentence after so many words. "X will do ABCDEFGHIJKLMNOPQRSTUVWXY and does Z."

I thought it could be a typo for "X will do ... and do Z" but only when I went to write why it's weird, I realised it doesn't have to be a typo.

It doesn't even matter here. It would mean the same thing either way.

Re: Everyone is building LLM routers, we deprecated ours

#38
post #18

I spent a lot of time researching LLM routing last year and also came to the conclusion that it's generally not worth the effort. It's too hard to understand the difficulty of a query a priori. One specific challenge I was seeing is that difficulty depends a lot on what information is retrievable by the agent. Consider the question "what is the 5-state busy beaver number?" ( https://en.wikipedia.org/wiki/Busy_beaver…

I mean can't you just directly test some variations of the query against many models at once and pick the cheapest model that hits your accuracy goal? If it's a one time run then ofc this is all pointless, but for ongoing tasks it makes sense. This seems more logical to me than making another AI model of some sorts intuit the right LLM for the job.

Re: Everyone is building LLM routers, we deprecated ours

#39

Some routing services not just route to different LLMs, they also handle all the legal issues (GDPR compliance, ISO certification, guaranteed Zero-Data-Retention, domestic data processing/European based clouds, etc.). In regulated industries, these things matter a lot, especially when processing of sensitive data is involved.

By "guaranteed" they mean "we promise really hard" right? Data retention has so much profit potential that any provider would have to be completely stupid not to retain data, even after promising they didn't.

Re: Everyone is building LLM routers, we deprecated ours

#40
post #18

I spent a lot of time researching LLM routing last year and also came to the conclusion that it's generally not worth the effort. It's too hard to understand the difficulty of a query a priori. One specific challenge I was seeing is that difficulty depends a lot on what information is retrievable by the agent. Consider the question "what is the 5-state busy beaver number?" ( https://en.wikipedia.org/wiki/Busy_beaver…

I think routing should be pushed 'down the stack' so to speak.

What I mean is that most tasks can be recursively fragmented into smaller tasks, and once you've hit suitable leaf nodes -- where the task is very granular -- you can begin to deterministically show which models perform better or worse for that specific task. Then, when your agent is running a workflow, you may use various models for different steps in a workflow. For example, some models may excel at exploration, some at determining a good architectural fit for an implementation, some at actually writing the implementation, and so on.

But you don't know until you define your 'work' taxonomy, and still further, you won't know until you have a statistically significant number of runs on a given chunk of work. Once you have that, though, you can hone in on models that excel at one specific task or another and prefer those the majority of the time (say ~80%) and hold back the remainder work as a 'test corpus' just in case a different or new model does even better.

This is something I've kept in the back of my head as I've been working through my agent harness primitives -- specifically enabling different models per chunk of work.

Post reply on HN