Co-founder and COO of OpenRouter here. Thanks everyone for the feedback here. Some of this we are aware of, some of it we aren't. Some we can fix, some of it is inherent to inference (and we in fact improve the situation dramatically). Philosophically, at OpenRouter we are trying to do two different things, that are sometimes at odds with one another: 1. Let you use a lot of capacity across a lot of providers, in a w…
So you want to use OpenRouter?
221–222 of 222 posts
Re: So you want to use OpenRouter?
#222Earlier quoted context omitted.
Thanks for the insight here. One thing to note about the first graph: nobody is doing as well as the first part on tool calling, and it's not close. This might be the fault of the other providers, but it's probably just something slightly different that the first party does with the model inference program than anybody else, and that's not sure to weights it's due to vLLM twiddling (or whatever) and probably becuase…
This matches what I found experimentally with gpt-oss-20b during OpenAI's red-teaming challenge. After moving from hosted inference to running the model myself on rented H100s via vast.ai, I saw the model refuse the same kinds of prompts at noticeably different rates depending on the inference stack — differences of roughly 5–10 percentage points with otherwise identical experimental parameters and seeds. So I very m…
Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way. (And if this was a false positive, I apologize!)