Viewing profile — lebovic
lebovic
HN member- Joined
- Tue, Apr 02, 2019, 3:42 PM UTC
- HN karma
- 1,931
- Public activity
- 143 items
- HN profile
- View on Hacker News ↗
About lebovic
Email is my firstname@lastname.com
Formerly at Anthropic, co-founder of Toolchest (a YC startup), and a bio startup
Recent public activity
-
comment
Comment #49186430
I formed a PBC and worked at a well-known PBC. Personally, I opted for a PBC because I liked that I could balance a specific cause with shareholder benefit. In most cases, it doesn…
-
comment
Comment #49051679
> all frontier models benefit from more tokens not just Kimi K3 Past a point, that doesn't hold and the score plateaus. Token hungry models tend to plateau at a much higher token c…
-
comment
Comment #49045371
The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...
-
comment
Comment #49045215
> Kimi K3 performs significantly below the most recent frontier cyber-capable models UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a toke…
-
comment
Comment #49031755
Almost exactly a year ago, we released Opus 4.1. It was definitely capable of finding vulnerabilities, and people were using custom harnesses to do so quite effectively. The newer …
-
comment
Comment #49031714
Yes, I think you could probably get something similar from Opus 4.5 (2025). Definitely Opus 4.6. I still think recent models are more capable, though! Some of the model behaviors t…
- comment
-
comment
Comment #48970644
(This comment was originally on another merged post, and "this page" referred to https://www.qwencloud.com/pricing/token-plan )
-
comment
Comment #48966205
I'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page. Looks like they're previewing the model only on their …
- story
-
comment
Comment #48963917
Kimi K3 only supports "max" reasoning effort right now, but they plan to enable other levels soon [1]. When I looked at traces from benchmarking, I saw a lot of backtracking and un…
- comment
- comment
- comment
- comment
- comment
-
comment
Comment #48749691
After spending years on a problem, it's exciting to see it start to get more attention and move towards being meaningfully solved. But I try to limit my time on HN, and I thought s…
- comment
-
comment
Comment #48748885
I assume they do hallucinate, just like with coding or finding vulnerabilities. You can try to minimize it (e.g. with a reviewer agent, which Claude Science and Biomni have), but n…
-
comment
Comment #48742310
I can't speak for Claude Science, but I prefer using Biomni as an agent for bio over Claude Code with a custom setup because a) Biomni stays on the frontier for bio, b) it has a co…
-
comment
Comment #48736916
I built one of the connected tools included in this launch (the Biomni HPC [1]), and I have spent an inordinate amount of my life working on this problem. (I also worked at Anthrop…
- story
-
comment
Comment #48714875
In this case, the benchmarks were private and it still outperformed.
-
comment
Comment #48713609
GLM 5.2 and DeepSeek v4 Pro seem to approach security research differently. This benchmark was with GLM 5.1, but the patterns are similar: https://dualuse.dev/posts/deepseek-v4-thi…
- comment