Viewing profile — mrifaki
mrifaki
HN member- Joined
- Tue, Apr 07, 2026, 3:10 AM UTC
- HN karma
- 2
- Public activity
- 3 items
- HN profile
- View on Hacker News ↗
About mrifaki
---
Reinforcement learning, mostly. The parts I keep circling back to are multi-agent learning, environment design, and the question of when an agent should stop gathering information and commit.
---
rifaki.me
---
mrifaki@stanford.edu
Recent public activity
-
comment
Comment #47802487
the adaptive thinking complaints in this thread are interesting because they are basically the same verifier quality problem showing up in a different costume the model has to deci…
-
comment
Comment #47747285
this is atctually he reward hacking problem from RL showing up in evaluation infra which is not surprising but worth naming clearly, an interesting question raised here is whether …
-
comment
Comment #47732568
finding vulns in a large codebase is a search problem with a huge negative space and what aisle measured is classification accuracy on ground-truth positives, those are different t…