Live data from Hacker News

Viewing profile — skysniper

skysniper

HN member
Joined
Fri, Mar 24, 2017, 5:26 AM UTC
HN karma
100
Public activity
30 items

About skysniper

No profile information was provided.

Recent public activity

  1. story
    Show HN: Inter-session messaging between Claude Code sessions

    I typically run multiple Claude Code sessions at the same time. Sometimes I need my existing sessions to work together, but CC does not support inter-session messaging so I have to…

  2. comment
    Comment #47798196

    Ran preliminary benchmarks on Opus 4.7, noticeably better than Opus 4.6, about 15% higher cost per task due to more tool calls, most performant and expensive model so far

  3. story
  4. comment
    Comment #47682385

    where are you Mythos

  5. story
  6. comment
    Comment #47608448

    check out my reply, his chart is plotting the wrong metric (average quality score)

  7. comment
    Comment #47607425

    i added native plot and stats for aggregated results, on arena page. please check it out!

  8. comment
    Comment #47607070

    yeah but i'm not using the free version for benchmark...

  9. comment
    Comment #47606949

    added https://app.uniclaw.ai/arena/model-stats also added per battle stats in battle detail page

  10. comment
    Comment #47606609

    I know, that was indeed a bad judge move. I've manually checked tens of tasks so far, and that one is one of the worst... I would say check a few more, judge has some noise but in …

  11. comment
    Comment #47606364

    well, I still want to use it but the first day i tried openclaw + opus, it costs me ~$500...

  12. comment
    Comment #47606236

    > The explanation is that network errors were credited with a quality score of 0, and there were _a lot_ of network errors. all network error, provider error, openclaw error are ex…

  13. comment
    Comment #47605473

    I will try and add it. But I doubt it works well because Mimo V2 Pro is beaten by stepfun even at performance leaderboard (price is not a factor in this leaderboard), so I expect M…

  14. comment
    Comment #47605413

    TBH that was my initial thought too, but I found some problem using this approach: Essentially I'm using the relative rank in each battle to fit a latent strength for each model, a…

  15. comment
    Comment #47605186

    it's actually pretty good at openclaw type of tasks for non technical users: lots of tool calls, some simple programing

  16. comment
    Comment #47604916

    what kind of tasks did you try?

  17. comment
    Comment #47604521

    I appreciate honest feedback, best way to learn :)

  18. comment
    Comment #47604478

    both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just wan…

  19. comment
    Comment #47604318

    I should have clarified I didn't use the free version...

  20. comment
  21. comment
    Comment #47604214

    > Is the judge an LLM? Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when comp…

  22. comment
    Comment #47603932

    another thing from the bench I didn't expect: gemini 3.1 pro is very unreliable at using skills. sometimes it just reads the skill and decide to do nothing, while opus/sonnet 4.6 a…

  23. comment
    Comment #47603694

    thanks for the info. before running the bench i only tried it in arena.ai type of tasks and it was not impressive. i didn't expect it to be that good at agentic tasks

  24. comment
    Comment #47603528

    all 300+ battle data are available at https://app.uniclaw.ai/arena/battles , every single battle is shown with raw conversional history, produced files, judge's verdict and final s…

  25. comment