Live data from Hacker News

Viewing profile — lebovic

lebovic

HN member
Joined
Tue, Apr 02, 2019, 3:42 PM UTC
HN karma
1,931
Public activity
143 items

About lebovic

Noah Lebovic

Email is my firstname@lastname.com

Formerly at Anthropic, co-founder of Toolchest (a YC startup), and a bio startup

Recent public activity

  1. comment
    Comment #49186430

    I formed a PBC and worked at a well-known PBC. Personally, I opted for a PBC because I liked that I could balance a specific cause with shareholder benefit. In most cases, it doesn…

  2. comment
    Comment #49051679

    > all frontier models benefit from more tokens not just Kimi K3 Past a point, that doesn't hold and the score plateaus. Token hungry models tend to plateau at a much higher token c…

  3. comment
    Comment #49045371

    The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...

  4. comment
    Comment #49045215

    > Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a toke…

  5. comment
    Comment #49031755

    Almost exactly a year ago, we released Opus 4.1. It was definitely capable of finding vulnerabilities, and people were using custom harnesses to do so quite effectively. The newer …

  6. comment
    Comment #49031714

    Yes, I think you could probably get something similar from Opus 4.5 (2025). Definitely Opus 4.6. I still think recent models are more capable, though! Some of the model behaviors t…

  7. comment
  8. comment
    Comment #48970644

    (This comment was originally on another merged post, and "this page" referred to https://www.qwencloud.com/pricing/token-plan )

  9. comment
    Comment #48966205

    I'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page. Looks like they're previewing the model only on their …

  10. story
  11. comment
    Comment #48963917

    Kimi K3 only supports "max" reasoning effort right now, but they plan to enable other levels soon [1]. When I looked at traces from benchmarking, I saw a lot of backtracking and un…

  12. comment
  13. comment
  14. comment
  15. comment
  16. comment
  17. comment
    Comment #48749691

    After spending years on a problem, it's exciting to see it start to get more attention and move towards being meaningfully solved. But I try to limit my time on HN, and I thought s…

  18. comment
  19. comment
    Comment #48748885

    I assume they do hallucinate, just like with coding or finding vulnerabilities. You can try to minimize it (e.g. with a reviewer agent, which Claude Science and Biomni have), but n…

  20. comment
    Comment #48742310

    I can't speak for Claude Science, but I prefer using Biomni as an agent for bio over Claude Code with a custom setup because a) Biomni stays on the frontier for bio, b) it has a co…

  21. comment
    Comment #48736916

    I built one of the connected tools included in this launch (the Biomni HPC [1]), and I have spent an inordinate amount of my life working on this problem. (I also worked at Anthrop…

  22. story
  23. comment
    Comment #48714875

    In this case, the benchmarks were private and it still outperformed.

  24. comment
    Comment #48713609

    GLM 5.2 and DeepSeek v4 Pro seem to approach security research differently. This benchmark was with GLM 5.1, but the patterns are similar: https://dualuse.dev/posts/deepseek-v4-thi…

  25. comment