Viewing profile — Trapais
Trapais
HN member- Joined
- Thu, Mar 23, 2023, 7:04 PM UTC
- HN karma
- 10
- Public activity
- 15 items
- HN profile
- View on Hacker News ↗
About Trapais
No profile information was provided.
Recent public activity
-
comment
Comment #41485676
Have you tried programming in something other than notepad.exe? Modern IDEs have auto complete so impressive, that you don't need to type much anyway. I type several symbols then s…
-
comment
Comment #40414424
For comparison, here's 8B [Nemotron]( https://huggingface.co/nvidia ): > 1,024 A100s were used for 19 days to train the model. > NVIDIA models are trained on a diverse set of publi…
-
comment
Comment #39298448
People say lots of stupid shit. If there is no code or even paper, there is no reason to believe. Beating benchmarks requires something more than a blind faith.
-
comment
Comment #39051585
>Making useful models is the goal. Sure, training datasets for pythia is useful. The Pile was used in lots of models. However it's hardly relevant that pythia itself was trained on…
-
comment
Comment #39043247
>Are there any true open-source LLM models, where all the training data is publicly-available (with a compatible license) Mamba has a version, trained on publicly available SlimPaj…
-
comment
Comment #39043090
OK. Where is your reproduction of Pythia trained from scratch? Or MPT? Or Amber? Shall we play a game where you give paper regarding pretraining (and we are not taling about puny m…
-
comment
Comment #39038964
You can grep for bad words. What you can't do(unless hoops are jumped through) is to verify that weights came from the same dataset. You can set the same random seed and still get …
-
comment
Comment #39023880
Propaganda. DoD already uses hollywood for propaganda: say nice things about Uncle Sam, let America save the day once again, and Uncle Sam will let you play with his toys. Now they…
-
comment
Comment #37748418
Looks like longformer to me. They just renamed "global attention" into "attention sink" and removed silly parts(distilled attention) and BERT parts([CLS] saw all N tokens, there is…
-
comment
Comment #36568879
> In my opinion, formed from over two decades of Linux, a piece of hardware having a libre driver written for it is the exact indicator of what can be relied upon to "just work". T…
-
comment
Comment #36381344
I have doubts it was extensively trained on German data. Who knows about GPT4, but GPT3 is ~92% of English and ~1.5% of German, which means it saw more "die, motherfucker, die" tha…
-
comment
Comment #36227541
That's a very fancy way with lots of fancy words to say "I have no idea how NN work, but if I sound smart maybe ppl will not figure out how stupid I sound". Well, you sound stupid …
-
comment
Comment #36192544
Maybe they should support their modern consumer cards in ROCm. Maybe their ROCm documentation should not suck balls. I'd say there is a reason AMD is a laughing stock in ML, but it…
-
comment
Comment #35279927
I will believe that it's open source the moment weights are downloaded on my computer and I don't need to summon DAN for using them: their goal is to make AI "safer", which is corp…
-
comment
Comment #35279732
[flagged]