Viewing profile — alexlitz
alexlitz
HN member- Joined
- Tue, Mar 19, 2019, 3:06 AM UTC
- HN karma
- 41
- Public activity
- 10 items
- HN profile
- View on Hacker News ↗
About alexlitz
No profile information was provided.
Recent public activity
- story
-
comment
Comment #47370316
The buried lede is this, if you have two dimensions and use rope, and hard-max attention you could simply store addresses as a given theta. With RoPE and sufficient precision that …
-
comment
Comment #47205068
> It baffles me that somebody capable of this kind of work would find this surprising. I should be clear I was not surprised that: 1) It struggled particularly hard with this sort …
-
comment
Comment #47197810
Yeah that is plausible enough.
-
comment
Comment #47191916
Yeah basically it is an implementation detail but most of them are zero, there is an equivalent 14 parameter sparse matrix for that.
-
comment
Comment #47189776
I imagine getting things to be polysemantic in a way that does not interfere would lead to sublinear scaling. Also there are smaller ones that were trained so would still be more l…
-
comment
Comment #47189547
For one the specific 36 parameter version is impossible without float64 so you might guess the corollary that it is not exactly amenable to being found by gradient descent. I think…
- story
-
comment
Comment #47189502
I made a blogpost on my submission (currently the top handwritten one at 36 parameters) https://alexlitzenberger.com/blog/building_a_minimal_transfo...
- story