Live data from Hacker News

Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

braw.dev

211–220 of 265 posts

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#211
post #109

I think you are kidding if you think you are going to be remotely approximately the quantity/quality of output you get from a $100/max sub with Zed/Openrouter. I easily get $1K+ of usage out of my $100 max sub. And that's with Opus 4.6 on high thinking.

> I easily get $1K+ of usage out of my $100 max sub. And that's with Opus 4.6 on high thinking. And people keep claiming the token providers are running inference at a profit.

In addition to usage distribution aspects others called out .

$1K is not actual cost, just API pricing being compared to subscription pricing. It is quite possible that API has a large operating margins, and say costs only $100 to deliver $1K worth of API credits.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#212
post #109

I think you are kidding if you think you are going to be remotely approximately the quantity/quality of output you get from a $100/max sub with Zed/Openrouter. I easily get $1K+ of usage out of my $100 max sub. And that's with Opus 4.6 on high thinking.

For personal use I've noticed Claude (via the web-based chat UI) making really bizarre mistakes lately like ignoring input or making completely random assumptions. At work Claude Code has turned into an absolute dog. It fails to follow instructions and builds stuff like a lazy junior developer without any architecture, tests, or verification. This is even with max effort, Opus 4.6, multiple agents, early compaction,…

Same, I'm looking hard for an alternative to what I had.

And I'm seeing the same thing in my sphere- everyone is bailing Anthropic the past few weeks. I figure that's why we're seeing more posts like this.

I hope they're paying attention.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#213
post #96
post #15

Has anyone (other than OpenClaw) used pi? ( https://shittycodingagent.ai/ , https://pi.dev/ ) Any insights / suggestions / best practices?

had been using claude max/opus with pi and the results have been incredible. Having pi write an AGENTS.md and dip your feet into creating your own skills specific to the project. With the anthropic billing change (not being able to use the max credits for pi) I think I have to cancel - as I'm whirring through credits now. Going to move to the $250/mo OpenAI codex plan for now.

Regardless of which harness you use, asking your agent to self-edit its own .claude (and to put it in the repo itself so you see the changes) is the single biggest impact change you can make in terms of compounding improvement. Couple this with telling it to create skills for /garden (clean up drift based on what changed this session), /handoff (garden, create ant skills to resolve friction encountered this session, and write a summary of the session and note for next agent), /takeover (read the latest handoff file). Since doing this I’ve completely cured my session-abandonment anxiety and can confidently swap to a new session at < 20% context usage without feeling like I’m talking to someone who just woke up from a coma.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#214
post #15

Has anyone (other than OpenClaw) used pi? ( https://shittycodingagent.ai/ , https://pi.dev/ ) Any insights / suggestions / best practices?

I bought a $30 Z.ai Coding Plan sub to go with it. 7 million tokens has only gone through 2% of my weekly usage using the GLM-5.1 model. I am pretty happy. I am only doing single project workflows, but with Z.ai I feel like it opens a whole new door to parallel workflows without hitting usage limits.

Is GLM-5.1 actually good?

I tested one of the other models that everyone is raving about yesterday (Qwen 3.6 plus) and within minutes found myself arguing with it even over a very simple task. After about 30 minutes (in which token usage never went over 50k because it was just me rewinding to give it more and more explicit instructions which it kept ignoring), I reverted everything and did it with Opus in literally about 4 minutes, after intentionally giving Opus a much more vague prompt.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#215
post #162
post #68

My 50c - ollama cloud 20$. GLM5 and kimi are really competitive models, Ollama usage limits insane high, no limits where to use (has normal APIs), privacy and no logging

Interesting. I've always been turned off by how vague the descriptions of Ollama's limits are for their paid tiers. What sort of work have you been doing with it?

Background agents (diy OpenClaw like), coding, assistant (openwebui).

The worst I saw - multiple parallel agents (opencode & pi-coding agents), with Kimi and glm, almost non stop development during the work day - 15-20% session consumption (I think it’s 2h bucket) max. Never hit the limit.

In contrast, 20$ Claude in the similar mode I consumed after just few hours of work.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#216
post #68

My 50c - ollama cloud 20$. GLM5 and kimi are really competitive models, Ollama usage limits insane high, no limits where to use (has normal APIs), privacy and no logging

yeah? why do you like that over using GLM5 in a VPS that charges by token use? $20 still cheaper and seamless to set up? how are the tokens per second?

I have roughly 20-40M token usage per day for GLM only (more if count other models). Using API pricing from OR it means ollama more profitable for me after day (few days if count cache properly).

For several models like Kimi and glm they have b300 and performance really good. At launch I got closer to 90-100 tps. Nowadays it’s around 60 tps stable across most models I used (utility models < 120B almost instant)

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#217

Earlier quoted context omitted.

One additional major benefit of OpenRouter is that there is no rate limiting. This is the primary reason why we went with OpenRouter because of the tight rate limiting with the native providers.

I think it's more accurate to say that they switch providers when there is rate limiting. The underlying provider can still limit rates. What Openrouter provides is automatic switching between providers for the same model. (I could be wrong.)

Beyond that, with some providers like Open AI, API limits are determined via a tiered account system based on your business relationship and spend.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#218
post #109

I think you are kidding if you think you are going to be remotely approximately the quantity/quality of output you get from a $100/max sub with Zed/Openrouter. I easily get $1K+ of usage out of my $100 max sub. And that's with Opus 4.6 on high thinking.

> I easily get $1K+ of usage out of my $100 max sub. And that's with Opus 4.6 on high thinking. And people keep claiming the token providers are running inference at a profit.

The model developers across the board stand by that most/all models are profitable by EOL, and losses come from R&D/Training.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#219

Earlier quoted context omitted.

I agree with you in certain circumstances, but not really for internal user inference. OpenRouter is great if you need to maintain uptime, but for basic usage (chat/coding/self-agents) you can do all of what you mentioned and more with a LiteLLM instance. The number of companies that send a bill is rarely a concern when it comes to “is work getting done”, but I agree with you that minimizing user friction is best. Fo…

The two things I like about OpenRouter: 1. The LLM provider doesn't know it's you (unless you have personally identifiable information in your queries). If N people are accessing GPT-5.x using OpenRouter, OpenAI can't distinguish the people. It doesn't know if 1 person made all those requests, or N. 2. The ability to ensure your traffic is routed only to providers that claim not to log your inputs (not even for secur…

1 - I can’t speak to whether that is the case with OpenRouter. However, I suspect that there is more than enough fingerprint and uniqueness inherent to the requests that an AI could probably do a fairly accurate job of reconstructing “possible” sources, even with such anonymity. The result is the same, all your information is still tied to OpenRouter in order to track the billing. That also ignores that OpenRouter is also privy to all that same information. In the end, it comes down to how much you trust your partners.

As for LiteLLM, the company you would pay for inference is going to know it is “you” — the account — but LiteLLM would also have the same effect of appearing to be a single source to that provider. That said, a uniqueness for a user may be passed (as is often with OpenRouter also) for security. Only you know who the users are, that never has to leave your network if you don’t want.

2 - well, you select the providers, so that’s pretty much on you? :-) basically, you are establishing accounts with the inference providers you trust. Bedrock has ZDR, SOC, HIPPA, etc available, even for token inference, as an example. Cost is higher without cache, but you can’t have true ZDR and Cache (that I know of), because a cache would have to be stored between requests. The closest you could get there is maybe a secure inference container but that piles on the cost. Still, plenty of providers with ZDR policies.

LiteLLM is effectively just a proxy for whatever supported (or OpenAI, Anthropic, etc compatible api provider) you choose.

Re: Reallocating $100/Month Claude Code Spend to Zed and OpenRouter

#220

Earlier quoted context omitted.

OpenRouter tells you if they submit with your user ID or anonymously if you hover over one of the icons on the provider, eg OpenAI has "OpenRouter submits API requests to this provider with an anonymous user ID.", Azure OpenAI on the other hand has "OpenRouter submits API requests to this provider anonymously.".

But does "anonymous user ID" mean that they make a user ID for you, and it's sticky? If I make a request today and another tomorrow, the same anonymous user ID is sent each time? Or do they keep changing it?

I believe they are static user ids that only OpenRouter knows is you (the anonymous part. Static id is required for any cached pricing. If the user id changes each request, it would be a massive security hole to reuse that cache between requests with different user ids.

Without caching, it would make sense to be per-request (more like a transaction-id, and would make sense to be) as this could then be tied internally back to a user while maintaining external anonymity, but unfortunately I don’t believe that is the case.

Post reply on HN