Live data from Hacker News

Viewing profile — xdotli

xdotli

HN member
Joined
Sat, Jul 08, 2023, 11:41 AM UTC
HN karma
16
Public activity
45 items

About xdotli

Founder BenchFlow.ai, a benchmark company.

Recent public activity

  1. story
  2. comment
    Comment #48275911

    How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3, Prime Intellect’s Intellect 3, Nous Research…

  3. story
  4. comment
    Comment #47693530

    Author here. We built 5 high-fidelity mock Google Workspace + Slack services and ran 7,224 trials across 6 frontier models and 4 agent harnesses. The headline finding that surprise…

  5. story
  6. comment
    Comment #47360904

    A two-week study of autonomous language model agents deployed in a live multi-party environment with persistent memory, email, shell access, and real human interaction — tested by …

  7. story
  8. comment
    Comment #47360639

    Agent scaffold comparison. We additionally evaluateOpenCode, an open-source scaffold that supports multiplemodel providers. Native CLI scaffolds consistently outper-form OpenCode w…

  9. story
  10. story
  11. story
  12. story
  13. story
  14. comment
    Comment #47119712

    yeah we didn't give agents access to the internet for creating their domain knowledge skills

  15. comment
    Comment #47119465

    The Register wrote about works on SkillsBench.ai

  16. story
  17. comment
    Comment #47058384

    no worries it's totally fine! there is indeed work needs to be done on the feedbacks generated skills. Thanks for helping us submitting on HackerNews. And for > a lot of Skills on …

  18. comment
    Comment #47055429

    20+ Anthropic Default Skills, 200k+ community skills on skillsmp. People talk about skills without knowing how well they work. We're hosting the largest Agent Skills hackathon at F…

  19. story
  20. comment
    Comment #47053442

    Did you check our repos and sites? the repo is skills native. Also please don't be misled by the original title, we have this configuration to eliminate the impact of internal know…

  21. comment
    Comment #47053218

    We collected 86 tasks from 105 domain experts across 11 domains, every task is verifiable, human created and has verified Skills. SOTA model without skills score ~30% without skill…

  22. story
  23. comment
    Comment #47052705

    we didn't create that headline yeah thanks for liking it

  24. comment
    Comment #47052689

    Thanks @dang for moderating! This is indeed not our original findings and this is a sub conclusion for an ablation we did to remove the confound of LLMs internal domain knowledge. …

  25. comment
    Comment #47052610

    I would frame the 'post-trajectory generated skills' as feedback-generated skills, so is Letta: https://www.letta.com/blog/skill-learning . We haven't seen existing research or hyp…