Live data from Hacker News

Viewing profile — gyang

gyang

HN member
Joined
Sat, Apr 22, 2017, 2:40 PM UTC
HN karma
127
Public activity
19 items

About gyang

No profile information was provided.

Recent public activity

  1. comment
    Comment #30991159

    I think the concept makes sense. The basic insight, that the right batch size depends on the difficulty and noisiness of a task, is already used by teams. For example, the PaLM pap…

  2. comment
    Comment #30990593

    I think there remains an immense amount of such suboptimality still hanging from the tree, so to speak. For example, our recent paper "Tensor Programs V: Tuning Large Neural Networ…

  3. story
  4. story
  5. story
  6. story
  7. story
  8. story
  9. comment
    Comment #21679965

    Right, so your argument would work if you allow the layer widths to tend to infinity sequentially (so this corresponds to finite networks where each previous layer is much bigger t…

  10. comment
    Comment #21668699

    Hi Radford, I'm very happy that this paper got your attention, but I'm regretful that I did not represent your research accurately! Indeed, in your thesis, section 2.3 is on "Prior…

  11. comment
    Comment #21660888

    I think in general, "training" a GP, i.e. doing GP inference (or kernel regression) is not done for speed reasons, but rather because they are sample efficient . More concretely, t…

  12. comment
    Comment #21660861

    Hi @zwaps, thanks so much for your interest! My answer to @throwlaplace seems relevant to your question, so let me copy it here and comment on your specific questions afterward. > …

  13. comment
    Comment #21660734

    @fgabriel mentioned this below: if the network is parametrized in a certain way, then the GP evolves according to a linear equation (if trained with square loss). In this linear eq…

  14. comment
    Comment #21654422

    Yes. I will have things to say about training, but that requires building up some theoretical foundations. This paper is the first step in laying it out. Stay tuned! :)

  15. comment
    Comment #21654419

    Hi, author here. Thanks for your interest! The CLT would be a good guess at approaching this problem, and indeed it is the approach of prior works [1][2]. But in this paper, the ke…

  16. comment
    Comment #21654327

    Hi, the author here. Thanks for your interest! Let me try answering some of your questions. > This sounds important and interesting but isn't wide the key word here? Yes, width is …

  17. story
  18. comment
    Comment #14172956

    Hi, I'm the lead author on this paper. If you have any questions please feel free to ask!

  19. story