I ran into a problem at work recently: we are given access to a bunch of models up to a full Claude Opus 4.8, but a monthly budget of 100k tokens. We are also given access to Gemini 3.5 Flash & 3.1 Pro with a daily budget of 50M tokens, but no tool calling. I'd love to hook Claude Code (or Pi) into the Gemini model, but the lack of tool-calling makes it quite difficult. I've been planning out how an intelligent route…
I’m curious how a workplace ends up with a model policy like this. It seems like you’d spend more time trying to work out how to use a tiny number of Opus tokens than doing it yourself.
Show HN: Smart model routing directly in Claude, Codex and Cursor
121–127 of 127 posts
Re: Show HN: Smart model routing directly in Claude, Codex and Cursor
#122Re: Show HN: Smart model routing directly in Claude, Codex and Cursor
#123Re: Show HN: Smart model routing directly in Claude, Codex and Cursor
#124Re: Show HN: Smart model routing directly in Claude, Codex and Cursor
#125Do you reward the RL model based on the token consumption when multiple LLMs complete the task ?
Re: Show HN: Smart model routing directly in Claude, Codex and Cursor
#126This is one of the most grotesque metric and funnels for data input I've seen.
And, to delete your Weave account? Email them, they don't respond.
Re: Show HN: Smart model routing directly in Claude, Codex and Cursor
#127> with no noticeable differences in quality or velocity. Have you done any A/B tests on this with evidence? (That's one thing I'd be very interested to see for claims like this - I'm not necessarily doubting you, it just seems like it could be useful to understand claims of quality/efficiency)
Great question! Our main product quantifies engineering productivity & quality so I think we're uniquely qualified to answer this - our velocity has only gone up and our quality (bugs introduced, code turnover) has not budged per our own analysis.