Live data from Hacker News

Viewing profile — mrifaki

mrifaki

HN member
Joined
Tue, Apr 07, 2026, 3:10 AM UTC
HN karma
2
Public activity
3 items

About mrifaki

Mouhssine Rifaki

---

Reinforcement learning, mostly. The parts I keep circling back to are multi-agent learning, environment design, and the question of when an agent should stop gathering information and commit.

---

rifaki.me

---

mrifaki@stanford.edu

Recent public activity

  1. comment
    Comment #47802487

    the adaptive thinking complaints in this thread are interesting because they are basically the same verifier quality problem showing up in a different costume the model has to deci…

  2. comment
    Comment #47747285

    this is atctually he reward hacking problem from RL showing up in evaluation infra which is not surprising but worth naming clearly, an interesting question raised here is whether …

  3. comment
    Comment #47732568

    finding vulns in a large codebase is a search problem with a huge negative space and what aisle measured is classification accuracy on ground-truth positives, those are different t…