Live data from Hacker News

Viewing profile — Turn_Trout

Turn_Trout

HN member
Joined
Mon, Dec 06, 2021, 8:23 PM UTC
HN karma
23
Public activity
16 items

About Turn_Trout

No profile information was provided.

Recent public activity

  1. comment
    Comment #49079412

    Those topics aren't on-topic for the essay. I've taken a pledge to donate at least 10% of my money to charity / impactful giving. A good sum of my donations have targeted high-impa…

  2. comment
    Comment #48928296

    Thank you for your praise. > I'd guess TurnTrout doesn't agree on that framing, otherwise he probably would not have been at Deep Mind. But clearly he and I agree on other ethical …

  3. story
  4. comment
    Comment #47691608

    I agree that they called many things remarkably well! That doesn't change the fact that AI 2027 is not a thing which happened, so it isn't valid to point out "this killed us in AI …

  5. comment
    Comment #47683624

    AI 2027 is not a real thing which happened. At best, it is informed speculation.

  6. story
  7. story
  8. comment
    Comment #44791856

    No one has empirically validated the so-called "most forbidden" descriptor. It's a theoretical worry which may or may not be correct. We should run experiments to find out.

  9. story
  10. comment
    Comment #44607871

    As someone who did their PhD in RL and alignment, it was not obvious to me a priori if, or when, or how badly obfuscation would be a problem. Yes, it's been predicted (and was pred…

  11. comment
    Comment #44380957

    > The #1 comment says that the rationality community is about "trying to reason about things from first principle", when if fact it is the opposite. Oh? Eliezer Yudkowsky (the most…

  12. story
  13. comment
    Comment #43178582

    They ran (at least) two control conditions. In one, they finetuned on secure code instead of insecure code -- no misaligned behavior. In the other, they finetuned on the same insec…

  14. comment
    Comment #39436215

    I'm the author of the GPT-2 work. This is a nice post, thanks for making it more available. :) Li et al[1] and I independently derived this technique last spring, and also someone …

  15. comment
    Comment #29476725

    First author here. Thanks for your comment! > there's a lot hidden in the "if physically possible" part of the quote from the paper: "Average-optimal agents would generally stop us…

  16. comment
    Comment #29469719

    Maybe you should read the paper, and/or the reviewer threads (as we discussed the nomenclature, and eventually agreed that "power" was accurate). We straightforwardly formalize a m…