Viewing profile — dpf
dpf
HN member- Joined
- Mon, Jul 01, 2013, 2:02 PM UTC
- HN karma
- 341
- Public activity
- 14 items
- HN profile
- View on Hacker News ↗
About dpf
No profile information was provided.
Recent public activity
- story
-
comment
Comment #35827595
code-davinci-002 is a base LM, and the other 3.5 models (text-davinci-{002,003}, gpt-3.5-turbo, and ChatGPT) use instruction tuning and/or RLHF. Source: https://platform.openai.com…
- story
- story
-
comment
Comment #16490253
Previous discussion: https://news.ycombinator.com/item?id=11951444 (2016) https://news.ycombinator.com/item?id=2591154 (2011)
- story
-
comment
Comment #14346665
These environments are often used as a testbed for reinforcement learning, e.g. https://arxiv.org/abs/1502.05477
- story
- story
- story
- story
- story
- story
- story