Live data from Hacker News

Viewing profile — dpf

dpf

HN member
Joined
Mon, Jul 01, 2013, 2:02 PM UTC
HN karma
341
Public activity
14 items

About dpf

No profile information was provided.

Recent public activity

  1. story
  2. comment
    Comment #35827595

    code-davinci-002 is a base LM, and the other 3.5 models (text-davinci-{002,003}, gpt-3.5-turbo, and ChatGPT) use instruction tuning and/or RLHF. Source: https://platform.openai.com…

  3. story
  4. story
  5. comment
    Comment #16490253

    Previous discussion: https://news.ycombinator.com/item?id=11951444 (2016) https://news.ycombinator.com/item?id=2591154 (2011)

  6. story
  7. comment
    Comment #14346665

    These environments are often used as a testbed for reinforcement learning, e.g. https://arxiv.org/abs/1502.05477

  8. story
  9. story
  10. story
  11. story
  12. story
  13. story
  14. story