Live data from Hacker News

Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

arxiv.org

11–20 of 29 posts

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#11
It seems they use 70% of the benchmark query-answer pairs to cluster and determine which models work best for each cluster (by sending all queries to all models and looking at responses vs ground truth answers). Then they route the remaining 30% "test" set queries according to those prior determinations. It doesn't seem surprising that this approach would give you Pareto efficiency on those benchmarks.

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#12

I’m fascinated by this new paradigm. We’ve more or less perfected Mixture-of-Experts inside a single model, where routing happens between subnetworks. What GPT-5 auto (and this paper) are doing is a step further: “LLM routing” across multiple distinct models. It’s still rough right now, but it feels inevitable that this will get much better over time.

Does this have a compute benefit or could one use different specialized LLM architectures / models for the subnetworks?

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#13

I’m fascinated by this new paradigm. We’ve more or less perfected Mixture-of-Experts inside a single model, where routing happens between subnetworks. What GPT-5 auto (and this paper) are doing is a step further: “LLM routing” across multiple distinct models. It’s still rough right now, but it feels inevitable that this will get much better over time.

I mean, agentic workflows have been a thing for a while now, this is just agentic chat.

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#16

It seems they use 70% of the benchmark query-answer pairs to cluster and determine which models work best for each cluster (by sending all queries to all models and looking at responses vs ground truth answers). Then they route the remaining 30% "test" set queries according to those prior determinations. It doesn't seem surprising that this approach would give you Pareto efficiency on those benchmarks.

It's ok if you can update the router over time, the more data you have the better.

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#17
Essentially, instead of modifying the prompt itself, the system intelligently directs the prompt to the LLM that is best suited to handle it based on its learned performance and efficiency characteristics for similar types of queries. It's externally optimizing people's prompts.

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#18
That's almost the most simple kind of router imaginable, isn't it? Just embed the query and route to the model that in the past has performed the best on similar queries?

I'm sure that has been documented/tried before, and this almost certainly doesn't work in practice. The typical counter-example would be to take a simple-sounding query that actually requires complex reasoning, but because the query is close in the embedding space to other simple-sounding queries, it would be sent to a "dumber model" for efficency.

I guess in their benchmarks that works out, because from what it sounds like, they do per-dataset clustering, so the embedding clusters may actually be able to cluster "complexity levels". However, if you were to mix all datasets into one (similar to how you would encounter it for most real-world use-cases) and cluster against that, this approach would surely break down.

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#19
post #15

Isn't this what NotDiamond (founded 2 years ago!) has been working to solve for? Maybe someone from their team will chime in (cc @t5-notdiamond)

Yeah that’s what my understanding is too about NotDiamomd. There are a bunch of similar products out there.

Re: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

#20

I’m fascinated by this new paradigm. We’ve more or less perfected Mixture-of-Experts inside a single model, where routing happens between subnetworks. What GPT-5 auto (and this paper) are doing is a step further: “LLM routing” across multiple distinct models. It’s still rough right now, but it feels inevitable that this will get much better over time.

I wish this could be exploited even further, where a big model could be built with a network of a lot of small, specialized, models

And then maybe you could just customize and optimize your own mode for local use. Almost like mixing and matching different modules. It would be nice to have a model that only knows and does what you need it to

Post reply on HN