Viewing profile — sjmaplesec
sjmaplesec
HN member- Joined
- Tue, Jul 31, 2018, 12:27 PM UTC
- HN karma
- 47
- Public activity
- 32 items
- HN profile
- View on Hacker News ↗
About sjmaplesec
No profile information was provided.
Recent public activity
-
comment
Comment #47850858
Full set of models tested: claude-opus-4-7 claude-opus-4-6 claude-sonnet-4-6 claude-haiku-4-5 gpt-5.4 gpt-5.3-codex gpt-5-codex cursor-composer-2 11 Skills used were here: https://…
- story
-
comment
Comment #47456632
Pass in a skill and it'll roast the contents: "This is not just useless - it is an insult to the very concept of functionality."
- story
-
comment
Comment #47340750
I ran some evals to see which Anthropic models use skills the best, between Opus, Sonnet and Haiku. I was pretty impressed how good Haiku was with skills at completing various task…
- story
-
comment
Comment #47268764
Link to all the review scans is here - mostly in the 50-70% range https://tessl.io/registry/skills/github/googleworkspace/cli
-
comment
Comment #47268738
There's so much more we can do around activation and skills creation. Looking at the eval results, there are even cases where the context makes the agent worse. Scenario 5, test 1 …
- story
-
comment
Comment #47166317
The review eval tests language, activation etc of skills. I guess you could move it all to a skill quick and then run an eval on that if using Tessl. This checks if the way you wri…
-
comment
Comment #47166294
An eval is to an LLM as a test is to code.
-
comment
Comment #47166285
Tessl can generate the evals, both to test anthropic best practices as well as running scenarios with and without the skill to check if it's helping
-
comment
Comment #47166263
Can add this as a skill or as part of a skill, and so you don't need to keep prompting the same things.
-
comment
Comment #47166257
No, the context can be human created as much as it could be llm generated. The suggestions are based on Anthropic best practices and allow the agents to activate, and use the skill…
- story
-
comment
Comment #46901143
This resonates with my experience: we have dozens of internal “playbooks” and prompt snippets floating around, and nobody knows which ones still work after model changes. If you ca…
-
comment
Comment #43547839
This is so true!
-
comment
Comment #43547766
I'm pretty sure Olivier Pomel rarely does podcasts, but this was a pretty good one. Some of my thoughts: - Customers "lie to themselves" saying they prefer noise to missed issues, …
- story
-
comment
Comment #41807171
Page title: Armon Dadgar, Hashicorp co-founder, on AI Native DevOps: Can AI shape the future of Autonomous DevOps workloads? Link is to an interesting podcast episode about Gen AI …
- story
- story
- story
- story
- story