Viewing profile — gregfrank
gregfrank
HN member- Joined
- Sat, Mar 21, 2026, 2:06 PM UTC
- HN karma
- -1
- Public activity
- 5 items
- HN profile
- View on Hacker News ↗
About gregfrank
AI alignment, mechanistic interpretability, model behavior, high-stakes agentic systems
Recent public activity
-
comment
Comment #47483164
[dead]
-
comment
Comment #47483150
[dead]
-
comment
Comment #47483143
"Trendslop" is a great name for something I think is a deeper structural problem than it appears. The issue isn't just that LLMs produce generic outputs, it's that our evaluation m…
-
comment
Comment #47483125
This framing points at something important that I think the alignment evaluation literature often misses: the distinction between what a model represents internally and what it doe…
-
comment
Comment #47483114
The "linear" assumption here is worth interrogating. In work I've been doing on alignment evaluation, I find that linear probes can achieve high accuracy on refusal-relevant direct…