Live data from Hacker News

Viewing profile — elexhobby

elexhobby

HN member
Joined
Thu, Dec 08, 2011, 3:23 PM UTC
HN karma
260
Public activity
12 items

About elexhobby

No profile information was provided.

Recent public activity

  1. comment
    Comment #36088617

    For another possibility, see https://arxiv.org/abs/2305.15717 . The new models may not actually be better - the evaluation may be broken.

  2. comment
    Comment #36080231

    GPT-4 is powerful over a diverse set of tasks. They use it to build a model which is better for a narrow sub-task. Pretty sure the model is sub-optimal to GPT-4 for everything else…

  3. comment
    Comment #36080055

    Not sure if this is obvious. But its incorrect to ditch on GPT-4. The paper uses self-instruct on GPT-4 to generate the training data on which it is fine-tuned. This paper would no…

  4. comment
    Comment #35984273

    This! The best resource I've found to explain transformers, that made them clear to me. I wish all deep learning papers were written like this, using pseudocode.

  5. comment
    Comment #33875912

    FWIW I've been told similar by Costco too, so Amazon isn't unique in this respect. My guess is that it adds some friction to the process. Refunding money is a purely digital activi…

  6. comment
    Comment #28613364

    Furthermore, follow https://twitter.com/_brohrer_/status/1425770502321283073 "When you have a problem, build two solutions - a deep Bayesian transformer running on multicloud Kuber…

  7. comment
    Comment #26656912

    any twitch streams you recommend?

  8. comment
    Comment #25703189

    Wow. I wish everything was explained so clearly. I understood everything in that post except this paragraph. """ Languages like Erlang must implement tail call optimizations, since…

  9. comment
    Comment #23505690

    Quote: "The election comes down to a few swing states, such as Pennsylvania, Wisconsin, and Michigan. Crucially, right now all three of those states have Democratic governors and R…

  10. story
  11. comment
    Comment #19510064

    Great post, thanks! Is there a reason why the training is started off with two separate matrices - the embedding and the context matrix? If the context matrix is anyway discarded a…

  12. story