Viewing profile — senseiV
senseiV
HN member- Joined
- Sat, Jul 01, 2023, 3:34 AM UTC
- HN karma
- 36
- Public activity
- 51 items
- HN profile
- View on Hacker News ↗
About senseiV
No profile information was provided.
Recent public activity
-
comment
Comment #40368287
if the ai is the product, and the product isnt trustable, isnt that a product issue??
-
comment
Comment #40252885
Does a TPU have XLA-graph for GPUs Cuda-graphs? Not sure on TPU theory
-
comment
Comment #40031343
Ive noticed the same on extremely small models aswell, magnitude is a positional encoding or a couple tokens, so its easy to grok?
-
comment
Comment #39828386
well world model in the context of the tulip fields, so models could be finetuned+sheared to drop size and remain effective
- story
-
comment
Comment #39635038
claude.ai
-
comment
Comment #39473695
Theres a startup doing that named galileo_ai
-
comment
Comment #39339840
Part of an FRC team building a Vision system from scratch, quite fun and nearly complete, just need to recalibrate some angle formulas
-
comment
Comment #39297624
I just saw a markdown mode show up today, but only partially, like bold and italics in markdown
-
comment
Comment #39295326
not sure if this is just chatgpt, but analogous evolution is interesting to see
-
comment
Comment #39229008
yes the size is different, but training a diffusion model and a language model are really different, like how RL models can be small but take a long time to train aswell
-
comment
Comment #39175921
Looking into the nordic pile maybe? There are some datasets
-
comment
Comment #39110472
NLP is not the industry, and a lot of research still goes into other things, like RL I've worked with several transformers competitors, and it def wont stay centralized on them
-
comment
Comment #39089344
GPT 2 and 3 used the p50K right? Then GPT-4 used cl100K
-
comment
Comment #39034535
> simulating entire AI-based societies. Didnt they already have scaled down simulations of this?
-
comment
Comment #38891798
replit/codesandbox maybe?
-
comment
Comment #38711076
? its better than GPT 2 for sure...
-
comment
Comment #38577287
V5 7b is out, close to hyena, gets 1400 t/s on a 3090, while an h100 llama 7b 8bit is 1200 t/s
-
comment
Comment #38577269
They Do, the latest rwkv v5, matches mamba at 3b scale, and from the benchmarks I see, its similar to hyena
-
comment
Comment #38486855
Just make a throwaway google?
-
comment
Comment #38455581
The so called "AI People" built the entire architecture, something people didn't think was possible at the scale and quality a year ago, and the matter of "artists should get whate…
-
comment
Comment #38439787
Bruh its simple physics, does one end or the other get lighter, by all measures we care about, not really, the mass of a proton or electron is beyond any consumer hardware measurem…
-
comment
Comment #38418027
He's talking about llama 2 superhots, and mistral derivatives that can be uncensored
-
comment
Comment #38389686
No, the orca 2 paper mentions more of a counter point towards NSFW and stuff, like if you gave it a NSFW prompt, it would retort back against it, which is arguably a good thing, bu…
-
comment
Comment #38381555
Where can you find those? I'm in the same situation as him, I've never heard of a 3d dataset better than objaverse XL. Got a public dataset?