Live data from Hacker News

Viewing profile — Trapais

Trapais

HN member
Joined
Thu, Mar 23, 2023, 7:04 PM UTC
HN karma
10
Public activity
15 items

About Trapais

No profile information was provided.

Recent public activity

  1. comment
    Comment #41485676

    Have you tried programming in something other than notepad.exe? Modern IDEs have auto complete so impressive, that you don't need to type much anyway. I type several symbols then s…

  2. comment
    Comment #40414424

    For comparison, here's 8B [Nemotron]( https://huggingface.co/nvidia ): > 1,024 A100s were used for 19 days to train the model. > NVIDIA models are trained on a diverse set of publi…

  3. comment
    Comment #39298448

    People say lots of stupid shit. If there is no code or even paper, there is no reason to believe. Beating benchmarks requires something more than a blind faith.

  4. comment
    Comment #39051585

    >Making useful models is the goal. Sure, training datasets for pythia is useful. The Pile was used in lots of models. However it's hardly relevant that pythia itself was trained on…

  5. comment
    Comment #39043247

    >Are there any true open-source LLM models, where all the training data is publicly-available (with a compatible license) Mamba has a version, trained on publicly available SlimPaj…

  6. comment
    Comment #39043090

    OK. Where is your reproduction of Pythia trained from scratch? Or MPT? Or Amber? Shall we play a game where you give paper regarding pretraining (and we are not taling about puny m…

  7. comment
    Comment #39038964

    You can grep for bad words. What you can't do(unless hoops are jumped through) is to verify that weights came from the same dataset. You can set the same random seed and still get …

  8. comment
    Comment #39023880

    Propaganda. DoD already uses hollywood for propaganda: say nice things about Uncle Sam, let America save the day once again, and Uncle Sam will let you play with his toys. Now they…

  9. comment
    Comment #37748418

    Looks like longformer to me. They just renamed "global attention" into "attention sink" and removed silly parts(distilled attention) and BERT parts([CLS] saw all N tokens, there is…

  10. comment
    Comment #36568879

    > In my opinion, formed from over two decades of Linux, a piece of hardware having a libre driver written for it is the exact indicator of what can be relied upon to "just work". T…

  11. comment
    Comment #36381344

    I have doubts it was extensively trained on German data. Who knows about GPT4, but GPT3 is ~92% of English and ~1.5% of German, which means it saw more "die, motherfucker, die" tha…

  12. comment
    Comment #36227541

    That's a very fancy way with lots of fancy words to say "I have no idea how NN work, but if I sound smart maybe ppl will not figure out how stupid I sound". Well, you sound stupid …

  13. comment
    Comment #36192544

    Maybe they should support their modern consumer cards in ROCm. Maybe their ROCm documentation should not suck balls. I'd say there is a reason AMD is a laughing stock in ML, but it…

  14. comment
    Comment #35279927

    I will believe that it's open source the moment weights are downloaded on my computer and I don't need to summon DAN for using them: their goal is to make AI "safer", which is corp…

  15. comment
    Comment #35279732

    [flagged]