Live data from Hacker News

Viewing profile — yohji1984

yohji1984

HN member
Joined
Fri, May 29, 2026, 1:21 AM UTC
HN karma
8
Public activity
15 items

About yohji1984

No profile information was provided.

Recent public activity

  1. story
  2. story
    Why My Open-Source Project Hasn't Done Better

    Since the beginning of 2026, many projects such as RTK, Caveman, and Ponytail have claimed that they can reduce token usage by 80–90%. Some of them gained tens of thousands of GitH…

  3. comment
    Comment #48969663

    [flagged]

  4. story
  5. comment
    Comment #48957523

    Sorry I will edit the post

  6. story
    Claude Code team should try macro so users can complete 3x as many tasks

    So the idea is really simple. If CC need to change a file and run some test, CC needs to: # Turn 1 — apply patch (change package file) # Turn 2 — apply patch (fix the bug) # Turn 3…

  7. story
  8. story
    Is GPT-5.6 Sol Max Worth It?

    I ran some test with gpt- 5.6 sol in max reasoning on both codex cli and my own agent harness: https://github.com/Tura-AI/tura I tested only 1 task and the toekn efficency differen…

  9. story
    Why people chasing after useless token saving plugins and ignoring real solution

    I wrote a blog yesterday on how useless RTK and Ponytail are on real coding tasks. And published my agent harness long-horizon task benchmarks on 80% real token saving. I just want…

  10. story
  11. story
  12. comment
    Comment #48922040

    Hi HN, I've been working on an agentic runtime framework. The earlier benchmark eval is promising. But I can see the limits of the test design and the fragility of the runtime itse…

  13. story
  14. comment
    Comment #48473477

    Sorry, I missed the Open Items section. You're right about that, designing a good eval harness can be difficult and expensive. Maybe we need some kind of community project for agen…

  15. comment
    Comment #48469285

    I'm wondering why all these token-saving solutions focus their benchmarks exclusively on simple Q&A tasks. If their tools truly saved money in real, long-term programming tasks, th…