Viewing profile — Turn_Trout
Turn_Trout
HN member- Joined
- Mon, Dec 06, 2021, 8:23 PM UTC
- HN karma
- 23
- Public activity
- 16 items
- HN profile
- View on Hacker News ↗
About Turn_Trout
No profile information was provided.
Recent public activity
-
comment
Comment #49079412
Those topics aren't on-topic for the essay. I've taken a pledge to donate at least 10% of my money to charity / impactful giving. A good sum of my donations have targeted high-impa…
-
comment
Comment #48928296
Thank you for your praise. > I'd guess TurnTrout doesn't agree on that framing, otherwise he probably would not have been at Deep Mind. But clearly he and I agree on other ethical …
- story
-
comment
Comment #47691608
I agree that they called many things remarkably well! That doesn't change the fact that AI 2027 is not a thing which happened, so it isn't valid to point out "this killed us in AI …
-
comment
Comment #47683624
AI 2027 is not a real thing which happened. At best, it is informed speculation.
- story
- story
-
comment
Comment #44791856
No one has empirically validated the so-called "most forbidden" descriptor. It's a theoretical worry which may or may not be correct. We should run experiments to find out.
- story
-
comment
Comment #44607871
As someone who did their PhD in RL and alignment, it was not obvious to me a priori if, or when, or how badly obfuscation would be a problem. Yes, it's been predicted (and was pred…
-
comment
Comment #44380957
> The #1 comment says that the rationality community is about "trying to reason about things from first principle", when if fact it is the opposite. Oh? Eliezer Yudkowsky (the most…
- story
-
comment
Comment #43178582
They ran (at least) two control conditions. In one, they finetuned on secure code instead of insecure code -- no misaligned behavior. In the other, they finetuned on the same insec…
-
comment
Comment #39436215
I'm the author of the GPT-2 work. This is a nice post, thanks for making it more available. :) Li et al[1] and I independently derived this technique last spring, and also someone …
-
comment
Comment #29476725
First author here. Thanks for your comment! > there's a lot hidden in the "if physically possible" part of the quote from the paper: "Average-optimal agents would generally stop us…
-
comment
Comment #29469719
Maybe you should read the paper, and/or the reviewer threads (as we discussed the nomenclature, and eventually agreed that "power" was accurate). We straightforwardly formalize a m…