Viewing profile — anayebi
anayebi
HN member- Joined
- Tue, Jun 02, 2026, 7:23 PM UTC
- HN karma
- 1
- Public activity
- 1 items
- HN profile
- View on Hacker News ↗
About anayebi
No profile information was provided.
Recent public activity
-
comment
Comment #48375852
No, the agents are not being adversarially prompted here. Rather, it's a consistent failure across models of RLHF-based safety-pretraining not generalizing to OOD open-ended agenti…