Live data from Hacker News

Viewing profile — senseiV

senseiV

HN member
Joined
Sat, Jul 01, 2023, 3:34 AM UTC
HN karma
36
Public activity
51 items

About senseiV

No profile information was provided.

Recent public activity

  1. comment
    Comment #40368287

    if the ai is the product, and the product isnt trustable, isnt that a product issue??

  2. comment
    Comment #40252885

    Does a TPU have XLA-graph for GPUs Cuda-graphs? Not sure on TPU theory

  3. comment
    Comment #40031343

    Ive noticed the same on extremely small models aswell, magnitude is a positional encoding or a couple tokens, so its easy to grok?

  4. comment
    Comment #39828386

    well world model in the context of the tulip fields, so models could be finetuned+sheared to drop size and remain effective

  5. story
  6. comment
    Comment #39635038

    claude.ai

  7. comment
    Comment #39473695

    Theres a startup doing that named galileo_ai

  8. comment
    Comment #39339840

    Part of an FRC team building a Vision system from scratch, quite fun and nearly complete, just need to recalibrate some angle formulas

  9. comment
    Comment #39297624

    I just saw a markdown mode show up today, but only partially, like bold and italics in markdown

  10. comment
    Comment #39295326

    not sure if this is just chatgpt, but analogous evolution is interesting to see

  11. comment
    Comment #39229008

    yes the size is different, but training a diffusion model and a language model are really different, like how RL models can be small but take a long time to train aswell

  12. comment
    Comment #39175921

    Looking into the nordic pile maybe? There are some datasets

  13. comment
    Comment #39110472

    NLP is not the industry, and a lot of research still goes into other things, like RL I've worked with several transformers competitors, and it def wont stay centralized on them

  14. comment
    Comment #39089344

    GPT 2 and 3 used the p50K right? Then GPT-4 used cl100K

  15. comment
    Comment #39034535

    > simulating entire AI-based societies. Didnt they already have scaled down simulations of this?

  16. comment
    Comment #38891798

    replit/codesandbox maybe?

  17. comment
    Comment #38711076

    ? its better than GPT 2 for sure...

  18. comment
    Comment #38577287

    V5 7b is out, close to hyena, gets 1400 t/s on a 3090, while an h100 llama 7b 8bit is 1200 t/s

  19. comment
    Comment #38577269

    They Do, the latest rwkv v5, matches mamba at 3b scale, and from the benchmarks I see, its similar to hyena

  20. comment
    Comment #38486855

    Just make a throwaway google?

  21. comment
    Comment #38455581

    The so called "AI People" built the entire architecture, something people didn't think was possible at the scale and quality a year ago, and the matter of "artists should get whate…

  22. comment
    Comment #38439787

    Bruh its simple physics, does one end or the other get lighter, by all measures we care about, not really, the mass of a proton or electron is beyond any consumer hardware measurem…

  23. comment
    Comment #38418027

    He's talking about llama 2 superhots, and mistral derivatives that can be uncensored

  24. comment
    Comment #38389686

    No, the orca 2 paper mentions more of a counter point towards NSFW and stuff, like if you gave it a NSFW prompt, it would retort back against it, which is arguably a good thing, bu…

  25. comment
    Comment #38381555

    Where can you find those? I'm in the same situation as him, I've never heard of a 3d dataset better than objaverse XL. Got a public dataset?