Live data from Hacker News

Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k

systima.ai

1–10 of 433 posts

Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k

#1
This started based off of a hunch. We usually use OpenCode, but were 'forced' to use Claude Code for a while due to issues with Meridian. In that time, we saw the usage meter rise much, much more quickly than when using OpenCode.

This was the initial anecdotal evidence, but we undertook this small study to collect empirical data:

We added logging between the agentic coding tool (Claude Code and OpenCode) and Anthropic's endpoint, and captured all requests (and the returned usage blocks).

With one caveat (toward the end of the post) we found unambiguously that Claude Code was far more inefficient in terms of its cache strategy and its harness token usage than OpenCode.

Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
systima.ai

Re: Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k

#2
> Claude Code 2.1.207 and OpenCode 1.17.18, both pinned to claude-sonnet-4-5

So not only is this article AI-written, but the testing was entirely done by AI, too? I can't see any other reason to use such an old model.

> Our traffic passes through a local LLM gateway that wraps requests in its own envelope, a constant we measured at roughly 6,200 tokens with bare calibration requests

Why do you need to do calibration requests to figure out how your own gateway is affecting requests?

> Its subagent lane did not complete cleanly through our gateway

> We attempted to toggle extended thinking in both harnesses and are declining to publish numbers. Our gateway applies its own thinking policy, neither harness's toggle demonstrably survived the path, and anything we quoted would be noise.

Why is your own gateway screwing with your testing?

Re: Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k

#5
This isn’t limited to large system prompts. Coding-agent harnesses are also becoming more aggressive about using tools, even for trivial requests. In our tests, prompts such as “Hey” or “commit” sometimes triggered 30+ tool calls:

https://quesma.com/blog/the-true-cost-of-saying-hi-to-an-ai-...

Tokenflation seems very real: the number of tokens consumed by simple tasks keeps increasing.

Re: Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k

#8
Early on in experimenting with local models, I found that hooking them up to Claude Code worked very well, but it was also really slow.

I used mitmproxy (setup assisted by Claude, natch) to capture Claude Code's entire initial system prompt and the whole thing was (I just double-checked) 162k of JSON.

This led me to start experimenting with Pi, OpenCode, and Hermes...

Re: Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k

#10
post #6

And pi agent is even less. The entire agent system prompt can be seen here: https://github.com/earendil-works/pi/blob/main/packages%2Fco...

If you really want a minimal agent that you heavily customize, just skip pi (130+ transitive dependencies on the "minimal" pi-coder package) and write your own. You learn a bunch, and it's not hard. You can even ask another LLM to help you get started.
Post reply on HN