Viewing profile — skysniper
skysniper
HN member- Joined
- Fri, Mar 24, 2017, 5:26 AM UTC
- HN karma
- 100
- Public activity
- 30 items
- HN profile
- View on Hacker News ↗
About skysniper
No profile information was provided.
Recent public activity
-
story
Show HN: Inter-session messaging between Claude Code sessions
I typically run multiple Claude Code sessions at the same time. Sometimes I need my existing sessions to work together, but CC does not support inter-session messaging so I have to…
-
comment
Comment #47798196
Ran preliminary benchmarks on Opus 4.7, noticeably better than Opus 4.6, about 15% higher cost per task due to more tool calls, most performant and expensive model so far
- story
-
comment
Comment #47682385
where are you Mythos
- story
-
comment
Comment #47608448
check out my reply, his chart is plotting the wrong metric (average quality score)
-
comment
Comment #47607425
i added native plot and stats for aggregated results, on arena page. please check it out!
-
comment
Comment #47607070
yeah but i'm not using the free version for benchmark...
-
comment
Comment #47606949
added https://app.uniclaw.ai/arena/model-stats also added per battle stats in battle detail page
-
comment
Comment #47606609
I know, that was indeed a bad judge move. I've manually checked tens of tasks so far, and that one is one of the worst... I would say check a few more, judge has some noise but in …
-
comment
Comment #47606364
well, I still want to use it but the first day i tried openclaw + opus, it costs me ~$500...
-
comment
Comment #47606236
> The explanation is that network errors were credited with a quality score of 0, and there were _a lot_ of network errors. all network error, provider error, openclaw error are ex…
-
comment
Comment #47605473
I will try and add it. But I doubt it works well because Mimo V2 Pro is beaten by stepfun even at performance leaderboard (price is not a factor in this leaderboard), so I expect M…
-
comment
Comment #47605413
TBH that was my initial thought too, but I found some problem using this approach: Essentially I'm using the relative rank in each battle to fit a latent strength for each model, a…
-
comment
Comment #47605186
it's actually pretty good at openclaw type of tasks for non technical users: lots of tool calls, some simple programing
-
comment
Comment #47604916
what kind of tasks did you try?
-
comment
Comment #47604521
I appreciate honest feedback, best way to learn :)
-
comment
Comment #47604478
both are shown in battle detail page already. Time is shown in Scores table. Number of tokens are shown in Cost details at the bottom of the Scores. (I thought most people just wan…
-
comment
Comment #47604318
I should have clarified I didn't use the free version...
- comment
-
comment
Comment #47604214
> Is the judge an LLM? Yes, judge is one of opus 4.6, gpt 5.4, gemini 3.1 pro (submitter can choose). Self judge (judge model is also one of the participants) is excluded when comp…
-
comment
Comment #47603932
another thing from the bench I didn't expect: gemini 3.1 pro is very unreliable at using skills. sometimes it just reads the skill and decide to do nothing, while opus/sonnet 4.6 a…
-
comment
Comment #47603694
thanks for the info. before running the bench i only tried it in arena.ai type of tasks and it was not impressive. i didn't expect it to be that good at agentic tasks
-
comment
Comment #47603528
all 300+ battle data are available at https://app.uniclaw.ai/arena/battles , every single battle is shown with raw conversional history, produced files, judge's verdict and final s…
- comment