Viewing profile — yaodub
yaodub
HN member- Joined
- Wed, May 20, 2026, 5:49 PM UTC
- HN karma
- 17
- Public activity
- 15 items
- HN profile
- View on Hacker News ↗
About yaodub
No profile information was provided.
Recent public activity
-
comment
Comment #48494338
[dead]
-
story
Show HN: Checkpoint! Airport security sim game built with Claude Fable
Fable just came out and I wanted to see what it was actually capable of. The game: an airport security sim where manage passengers and keep them from getting angry. You place rope …
-
comment
Comment #48490782
Your coworkers haven't changed. What changed is that people can hand off work they never had to think through themselves. So you don't know what they checked and you don't know wha…
-
comment
Comment #48483689
[dead]
-
comment
Comment #48480643
Sure, but you're paying the higher rate just to find out if that's true for your use case.
-
comment
Comment #48480600
On capability, yeah. On pricing, they just told most developers they're not the target anymore.
-
comment
Comment #48464498
SWE-Bench measures single tasks in isolation. In a real loop the model usually loses track of what I was trying to do long before code quality becomes the issue.
-
comment
Comment #48464405
Depends whether "unable to fully automate" means "needs occasional human checkpoints" or "slowly stops caring about your actual goal." Pretty different.
-
comment
Comment #48464321
[dead]
-
comment
Comment #48451348
Built a quant system that reads earnings transcripts for what management is trying not to say. The model is surprisingly bad at this. Turns out management is too.
-
comment
Comment #48431436
Interesting middle ground between full WASM lockdown and a bare environment. Did you end up needing to block anything else beyond the package manager?
-
comment
Comment #48428505
Same with legal questions, tbh. Spots the same issues, completely disagree on which ones matter. Maybe you just need a third model to choose between the outputs lol
- comment
-
comment
Comment #48420815
Have you found integrating outputs from different frontier labs consistently improves final results, or is it just kind of voodoo?
-
comment
Comment #48375701
[dead]