Viewing profile — elexhobby
elexhobby
HN member- Joined
- Thu, Dec 08, 2011, 3:23 PM UTC
- HN karma
- 260
- Public activity
- 12 items
- HN profile
- View on Hacker News ↗
About elexhobby
No profile information was provided.
Recent public activity
-
comment
Comment #36088617
For another possibility, see https://arxiv.org/abs/2305.15717 . The new models may not actually be better - the evaluation may be broken.
-
comment
Comment #36080231
GPT-4 is powerful over a diverse set of tasks. They use it to build a model which is better for a narrow sub-task. Pretty sure the model is sub-optimal to GPT-4 for everything else…
-
comment
Comment #36080055
Not sure if this is obvious. But its incorrect to ditch on GPT-4. The paper uses self-instruct on GPT-4 to generate the training data on which it is fine-tuned. This paper would no…
-
comment
Comment #35984273
This! The best resource I've found to explain transformers, that made them clear to me. I wish all deep learning papers were written like this, using pseudocode.
-
comment
Comment #33875912
FWIW I've been told similar by Costco too, so Amazon isn't unique in this respect. My guess is that it adds some friction to the process. Refunding money is a purely digital activi…
-
comment
Comment #28613364
Furthermore, follow https://twitter.com/_brohrer_/status/1425770502321283073 "When you have a problem, build two solutions - a deep Bayesian transformer running on multicloud Kuber…
-
comment
Comment #26656912
any twitch streams you recommend?
-
comment
Comment #25703189
Wow. I wish everything was explained so clearly. I understood everything in that post except this paragraph. """ Languages like Erlang must implement tail call optimizations, since…
-
comment
Comment #23505690
Quote: "The election comes down to a few swing states, such as Pennsylvania, Wisconsin, and Michigan. Crucially, right now all three of those states have Democratic governors and R…
- story
-
comment
Comment #19510064
Great post, thanks! Is there a reason why the training is started off with two separate matrices - the embedding and the context matrix? If the context matrix is anyway discarded a…
- story