Viewing profile — octbash
octbash
HN member- Joined
- Thu, Oct 31, 2019, 6:17 PM UTC
- HN karma
- 36
- Public activity
- 18 items
- HN profile
- View on Hacker News ↗
About octbash
No profile information was provided.
Recent public activity
-
comment
Comment #23607581
>I found this part to be the most interesting. If this current proclamation is a prelude to the overhaul of H1B system in a way that would make it work like described above, then i…
-
comment
Comment #23127374
I seem to find myself in the minority, but I don't think distill.pub is a particularly ideal model for publicizing research. distill.pub heavily favors fancy and interactive visual…
-
comment
Comment #22648461
How is that draconian? If you're issued a stay-home notice, STAY AT HOME.
-
comment
Comment #22648119
Excuse me - what "draconian quarantine laws"?
-
comment
Comment #22294548
One is a language generation model, the other is a fill-in-the-blank model. It sounds like they might be similar, but in practice they are different enough objectives (and in parti…
-
comment
Comment #22294535
Those are question-answering and language-understanding benchmarks respectively, neither of which has been suitable for language generation mode evaluation since GPT-1 was roundly …
-
comment
Comment #22199515
Is that essentially repeating the position embedding? I'm surprised that works, since the model should have no way to distinguish between the (e.g.) 1st and 513th token. (If I'm un…
-
comment
Comment #22197774
How are you using GPT-2 with an expanded context window? I was under the impression that the maximum context window was fixed.
-
comment
Comment #22197738
My counter-arguments (as a huge PyTorch fan) are: 1. GPT hasn't really been about model/architectural experimentation, just scale. GPT-2 and GPT were architecturally very similar. …
-
comment
Comment #22194464
Minimum=/=lowest effort. Especially with the caveat of "with a programming background", it is far easier to reason and debug through PyTorch with just Python knowledge, compared to…
-
comment
Comment #22127806
Besides vocabulary/characters, in what sense is Mandarin not simple?
-
comment
Comment #22097263
Yes, the Reformer is basically trading off noisier for faster training / memory savings.
-
comment
Comment #22095890
You keep insisting that they're the same when they're not, and then you try to subtly expand your original claim of "using less memory by hashing" to "to be more memory and compute…
-
comment
Comment #21993029
This sounds like some sort of perverse inverse of "This is good for bitcoin." Making progress of NLP benchmarks? Must be a sign that we're moving even closer to an even longer and …
-
comment
Comment #21889362
For the full infuriating timeline: https://landmarkscomplaints.com/chronology-of-events/
-
comment
Comment #21734606
How is it like linear regression, considering linear regression has a closed-form solution, and even if you're using an iterative solver, is a convex problem.
-
comment
Comment #21622814
The quick answer to your broader question is that there are multiple ways of slicing how GDP is computed, and yes, economists are aware of them and have thought through the edge-ca…
-
comment
Comment #21537411
> This and other discussions on HN make me think on the fact that so much of your voice has been recorded by companies when you call their support lines. Are these troves of voice …