Live data from Hacker News

Adaptive LLM routing under budget constraints

arxiv.org

71–80 of 83 posts

Re: Adaptive LLM routing under budget constraints

#72

Earlier quoted context omitted.

I mean I wish there were a plateau, without one we're well onto our way into techno-feudalism. I just don't see it.

That's what it is: wishful thinking. A lot of people really, really want AI tech to fail - because they don't like the alternative.

Yeah, obviously nobody that actually though about the consequences wants a large part of the population to become unemployed. Even if your job is not threatened by automation, it will be threatened by a lot of people looking for new jobs.

And the kind of automation brought by LLMs is decidely different than automation in the past which almost always created new (usually better) jobs. LLMs won't do this (at least to extent where it would matter) I think. Most people in ten years will have worse jobs (more physically straining, longer hours, less pay) unless there will be a political intervention.

Re: Adaptive LLM routing under budget constraints

#73
post #35

Earlier quoted context omitted.

It's trivial to get better score than GPT-4 with 1% of the cost by using my propertiary routing algorithm that routes all requests to Gemini 2.5 Flash. It's called GASP (Gemini Always, Save Pennies)

Does anyone working in an individual capacity actually end up paying for Gemini (Flash or Pro)? Or does Google boil you like a frog and you end up subscribing?

If I actually had time to work on my hobby projects Gemini pro would be the first thing I’d spend money on. As is, it’s amazing how much progress you can squeeze out of those 5 chats every 24h; I can get a couple hours of before-times hacking done in 15 minutes, which is incidentally when free usage gets throttled and my free time runs out.

Re: Adaptive LLM routing under budget constraints

#74

I'm very curious whether a) anecdotally, anyone has encountered a real enterprise cost-cutting effort focused on LLM APIs and b) empirically, whether anyone has done any research on price elasticity in LLMs of different performance scales. So far, my experience has been that it's just too early for most people / applications to worry about cost - at most, I've seen AI to be accountable for 10% of cloud costs. But ver…

In which context? The serious engineering folks are still in exploration phase, costs is mostly not a concern as long as shipping velocity increases. Reselling repackaged tokens is a different beast, no experience here

Re: Adaptive LLM routing under budget constraints

#75

Is this really the frontier of LLM research? I guess we really aren't getting AGI any time soon, then. It makes me a little less worried about the future, honestly. Edit: I never actually expected AGI from LLMs. That was snark. I just think it's notable that the fundamental gains in LLM performance seem to have dried up.

First, I don't think we will ever get to AGI. Not because we won't see huge advances still, but AGI is a moving ambiguous target that we won't get consensus on. But why does this paper impact your thinking on it? It is about budget and recognizing that different LLMs have different cost structures. It's not really an attempt to improve LLM performance measured absolutely.

Given OpenAI definition I’d expect AGI to be around in a decade or two. I don’t expect skynet, though maybe it’s a more realistic vision outcome that just droids mixing with humans.

Re: Adaptive LLM routing under budget constraints

#76
post #35

Earlier quoted context omitted.

It's trivial to get better score than GPT-4 with 1% of the cost by using my propertiary routing algorithm that routes all requests to Gemini 2.5 Flash. It's called GASP (Gemini Always, Save Pennies)

Does anyone working in an individual capacity actually end up paying for Gemini (Flash or Pro)? Or does Google boil you like a frog and you end up subscribing?

You get 1500 prompts on AIStudio across a few Gemini flash models. I think I saw 250 or 500 for 2.5. It’s basically free and beats the consumer rate limits of big apps (Claude, ChatGPT, Gemini, meta). I wonder when they’ll cut this off.

Re: Adaptive LLM routing under budget constraints

#77
post #2

Is there a reason human preference data is even needed? Don't LLMs already have a strong enough notion of question complexity to build a dataset for routing?

This is like asking someone to make you a sandwich and expect them to read your mind to determine what kind of sandwich you want.

Re: Adaptive LLM routing under budget constraints

#78

I'm very curious whether a) anecdotally, anyone has encountered a real enterprise cost-cutting effort focused on LLM APIs and b) empirically, whether anyone has done any research on price elasticity in LLMs of different performance scales. So far, my experience has been that it's just too early for most people / applications to worry about cost - at most, I've seen AI to be accountable for 10% of cloud costs. But ver…

LLM is far from the highest AI related cost, so we basically don't care about optimizing LLMs.

Obviously we don't use the super expensive ones like GPT4.5 or so. But we don't really bother with mini models, because GPT4.1 etc.. are cheap enough.

Stuff like speech to text etc.. are still way more expensive, and yes there we do focus on cost optimization. We have no large scale image generation use cases (yet)

Re: Adaptive LLM routing under budget constraints

#79
post #30

Earlier quoted context omitted.

So you don't expect AGI to be possible ever? Or is your concern mainly with the wildly different definitions people use for it and that we'll continue moving goal posts rather than agree we got there?

There's no concrete evidence AGI is possible mostly because it has no concrete definition. It's mostly hand waving, hype and credulity, and unproven claims of scalability right now. You can't move the goal posts because they don't exist.

even AI does not have a concrete definition.

Doesn't mean there aren't practical definitions depending on the context.

In essence, teaching an AI using recources meant for humans, and nothing more, would be considered AGI. That could be a practical definition, without needing much more rigour.

There is indeed no evidence we'll get there. But there is also no evidence LLM's should work as well as they do

Re: Adaptive LLM routing under budget constraints

#80

Is this really the frontier of LLM research? I guess we really aren't getting AGI any time soon, then. It makes me a little less worried about the future, honestly. Edit: I never actually expected AGI from LLMs. That was snark. I just think it's notable that the fundamental gains in LLM performance seem to have dried up.

just because it’s on arxiv doesn’t mean anything arxiv is essentially a blog under an academic format, popular amongst asian and south asian academic communities currently you can launder reputation with it, just like “white papers” in the crypto world allowed for capital for some time this ability will diminish as more people catch on

arxiv should really have a big red banner "NOT REVIEWED - DON'T USE AS A SOURCE" or something
Post reply on HN