Live data from Hacker News

Viewing profile — gregfrank

gregfrank

HN member
Joined
Sat, Mar 21, 2026, 2:06 PM UTC
HN karma
-1
Public activity
5 items

About gregfrank

AI alignment, mechanistic interpretability, model behavior, high-stakes agentic systems

Recent public activity

  1. comment
  2. comment
  3. comment
    Comment #47483143

    "Trendslop" is a great name for something I think is a deeper structural problem than it appears. The issue isn't just that LLMs produce generic outputs, it's that our evaluation m…

  4. comment
    Comment #47483125

    This framing points at something important that I think the alignment evaluation literature often misses: the distinction between what a model represents internally and what it doe…

  5. comment
    Comment #47483114

    The "linear" assumption here is worth interrogating. In work I've been doing on alignment evaluation, I find that linear probes can achieve high accuracy on refusal-relevant direct…