Viewing profile — 0xkvyb
0xkvyb
HN member- Joined
- Sat, May 02, 2026, 5:26 PM UTC
- HN karma
- 24
- Public activity
- 19 items
- HN profile
- View on Hacker News ↗
About 0xkvyb
No profile information was provided.
Recent public activity
-
comment
Comment #48492297
I’m experimenting with a small self-harness repo based on this paper. The idea is to run simulated users through an agent harness, collect the traces, group the recurring failures,…
- story
-
comment
Comment #48140709
The home computer is finally obsolete!
- story
-
comment
Comment #48140493
it’s crazy, we’re at a point where I commit code I haven’t seen, reviewed by another AI, followed up to by another AI and it’s just kind of scary. This thing will explode in our fa…
-
comment
Comment #48140428
but how would you do that? what about homework and coursework? students will just transcribe claude slop on paper and submit that.
-
comment
Comment #48140336
I was thinking recently that we might come to a point where we will at most write pseudo code. Face the facts, LLMs are pretty stellar at writing code, so why compete with them?
-
comment
Comment #48140287
If it really is, this would be a game changer. Gemini 3 flash is already very good, and is the hidden workhorse making many agentic routines possible. An upgrade that could reach f…
- story
-
comment
Comment #48140230
I think that universities just have to adapt to deal with slop, or think of new ways to challenge people to learn the essence of their studies. I wouldn’t want to be a uni teacher …
- comment
- comment
- comment
-
comment
Comment #47993969
Totally agree with you. There is only so much time before SF tech runs out of subsidy bucks, and Chinese models take the consumer spotlight
-
comment
Comment #47993142
[dead]
-
comment
Comment #47993129
[dead]
-
comment
Comment #47988579
Yes, GLM 5.1 is surprisingly good! Particularly for long-horizon Agentic tasks, with 100+ available tools. It really shocked me in a good way when it was able to complete a long ru…
-
comment
Comment #47988550
It might be at the frontier, but DeepSeek is really struggling with compute. The amount of 429 Rate Limit responses I've been getting just testing this thing made me pause all my a…
-
comment
Comment #47988505
Still crazy how easy it is to "jailbreak" even SOTA LLMs with a simple assistantResponse replacement in chat thread.