Viewing profile — pico_creator
pico_creator
HN member- Joined
- Thu, Mar 24, 2022, 3:17 PM UTC
- HN karma
- 268
- Public activity
- 99 items
- HN profile
- View on Hacker News ↗
About pico_creator
Recent public activity
-
comment
Comment #43550746
(original article author) I view it more as a shortcut. We have trained 7B and 14B models from scratch, matching transformer performance with similar sized datasets. This has been …
- story
-
comment
Comment #42573888
There is work done for Vision RWKV, and audio RWKV, an example paper is here: https://arxiv.org/abs/2403.02308 Its the same principle as open transformer models where an adapter is…
-
comment
Comment #42573878
One of the interesting "new direction" for RWKV and Mamba (or any recurrent model), is the monitoring and manipulation of the state in between token. For steerability, alignment, e…
-
comment
Comment #42573863
Not sure how indepth you want it to be. But we did do a co-presentation with one of the coauthors of mamba at latent space : https://www.youtube.com/watch?v=LPe6iC73lrc
-
comment
Comment #42573857
There is a current lack of "O1 style" reasoning dataset in open source space. QWQ did not release their dataset. So that would take some time for the community to prepare. It's def…
-
comment
Comment #42573850
kinda on a todo list, the model is open source on HF for anyone who is willing to make it work with lmarena
-
comment
Comment #42573843
lower compute cost especially over longer sequence length. Depending on context length, its 10x, 100x, or even 1000x+ cheaper. (quadratic vs linear cost difference)
-
comment
Comment #42573835
RWKV already solve the parallel compute problem for GPU, based on the changes it has done - so it is a recurrent model that can scale to thousands++ of GPU no issue. Arguably with …
-
comment
Comment #42572479
Currently the strongest RWKV model is 32B in size: https://substack.recursal.ai/p/q-rwkv-6-32b-instruct-preview This is a full drop in replacement for any transformer model use cas…
-
comment
Comment #42572474
Hey there, im Eugene / PicoCreator - co-leading the RWKV project - feel free to AMA =)
-
comment
Comment #42572466
This is actually the hypothesis for cartesia (state space team), and hence their deep focus on voice model specifically. Taking full advantage of recurrent models constant time com…
-
comment
Comment #42572451
Not an MoE, but we have already done hybrid models. And found it to be highly performant (as per the training budget) https://arxiv.org/abs/2407.12077
-
comment
Comment #41807545
Someone is losing the money. It’s elaborated in the article how and why this happens TLDR, VC money, is being burnt/lost
-
comment
Comment #41807232
Im quite sure there is more than a 100 clusters even. Though that would be harder to prove. So yea, it would be rough
-
comment
Comment #41807123
I actually signed up for separate new account, to double check that my business account was not being favored or rigged in "private beta" Its really not that hard to validate this …
-
comment
Comment #41807100
Not at $0.5 (which the lower bound in their marketing), but $1.5 is very doable on right times (done so multiple times) The article says $2. Which is quite consistent for a small c…
-
comment
Comment #41807087
Yup, but they at-least know where all these "small unused clusters" are. Bag holders, do not want to be shouting to the world they are bag holders.
-
comment
Comment #41806914
Also: how many of those consultants, have actually rented GPU's - used them for inference - or used them to finetune / train
-
comment
Comment #41806897
Do we have actual fp8 numbers? (or i could proxy it by /2 the fp4)
-
comment
Comment #41806863
Feel free to forward to the clients of "paid consultant". Also how do i collect my cut.
-
comment
Comment #41806765
Given their rising stock price trend, due to their moves in AI. Definitely worth it for them
-
comment
Comment #41806746
I really suggest shopping around. <$2 SXM is a real thing, if your patient enough on the schedule.
-
comment
Comment #41806736
Makes sense, though only folks like runpod / sfcompute / etc, have enough visibility to maybe pull this off? Its a risker move - then just taxing the excess compute now, and print …
-
comment
Comment #41806715
Only if ur a collector (so no if ur plugging it in)