Live data from Hacker News

Viewing profile — pico_creator

pico_creator

HN member
Joined
Thu, Mar 24, 2022, 3:17 PM UTC
HN karma
268
Public activity
99 items

About pico_creator

@picocreator (github / twitter)

Recent public activity

  1. comment
    Comment #43550746

    (original article author) I view it more as a shortcut. We have trained 7B and 14B models from scratch, matching transformer performance with similar sized datasets. This has been …

  2. story
  3. comment
    Comment #42573888

    There is work done for Vision RWKV, and audio RWKV, an example paper is here: https://arxiv.org/abs/2403.02308 Its the same principle as open transformer models where an adapter is…

  4. comment
    Comment #42573878

    One of the interesting "new direction" for RWKV and Mamba (or any recurrent model), is the monitoring and manipulation of the state in between token. For steerability, alignment, e…

  5. comment
    Comment #42573863

    Not sure how indepth you want it to be. But we did do a co-presentation with one of the coauthors of mamba at latent space : https://www.youtube.com/watch?v=LPe6iC73lrc

  6. comment
    Comment #42573857

    There is a current lack of "O1 style" reasoning dataset in open source space. QWQ did not release their dataset. So that would take some time for the community to prepare. It's def…

  7. comment
    Comment #42573850

    kinda on a todo list, the model is open source on HF for anyone who is willing to make it work with lmarena

  8. comment
    Comment #42573843

    lower compute cost especially over longer sequence length. Depending on context length, its 10x, 100x, or even 1000x+ cheaper. (quadratic vs linear cost difference)

  9. comment
    Comment #42573835

    RWKV already solve the parallel compute problem for GPU, based on the changes it has done - so it is a recurrent model that can scale to thousands++ of GPU no issue. Arguably with …

  10. comment
    Comment #42572479

    Currently the strongest RWKV model is 32B in size: https://substack.recursal.ai/p/q-rwkv-6-32b-instruct-preview This is a full drop in replacement for any transformer model use cas…

  11. comment
    Comment #42572474

    Hey there, im Eugene / PicoCreator - co-leading the RWKV project - feel free to AMA =)

  12. comment
    Comment #42572466

    This is actually the hypothesis for cartesia (state space team), and hence their deep focus on voice model specifically. Taking full advantage of recurrent models constant time com…

  13. comment
    Comment #42572451

    Not an MoE, but we have already done hybrid models. And found it to be highly performant (as per the training budget) https://arxiv.org/abs/2407.12077

  14. comment
    Comment #41807545

    Someone is losing the money. It’s elaborated in the article how and why this happens TLDR, VC money, is being burnt/lost

  15. comment
    Comment #41807232

    Im quite sure there is more than a 100 clusters even. Though that would be harder to prove. So yea, it would be rough

  16. comment
    Comment #41807123

    I actually signed up for separate new account, to double check that my business account was not being favored or rigged in "private beta" Its really not that hard to validate this …

  17. comment
    Comment #41807100

    Not at $0.5 (which the lower bound in their marketing), but $1.5 is very doable on right times (done so multiple times) The article says $2. Which is quite consistent for a small c…

  18. comment
    Comment #41807087

    Yup, but they at-least know where all these "small unused clusters" are. Bag holders, do not want to be shouting to the world they are bag holders.

  19. comment
    Comment #41806914

    Also: how many of those consultants, have actually rented GPU's - used them for inference - or used them to finetune / train

  20. comment
    Comment #41806897

    Do we have actual fp8 numbers? (or i could proxy it by /2 the fp4)

  21. comment
    Comment #41806863

    Feel free to forward to the clients of "paid consultant". Also how do i collect my cut.

  22. comment
    Comment #41806765

    Given their rising stock price trend, due to their moves in AI. Definitely worth it for them

  23. comment
    Comment #41806746

    I really suggest shopping around. <$2 SXM is a real thing, if your patient enough on the schedule.

  24. comment
    Comment #41806736

    Makes sense, though only folks like runpod / sfcompute / etc, have enough visibility to maybe pull this off? Its a risker move - then just taxing the excess compute now, and print …

  25. comment
    Comment #41806715

    Only if ur a collector (so no if ur plugging it in)