Viewing profile — desideratum
desideratum
HN member- Joined
- Wed, Jul 24, 2019, 9:24 PM UTC
- HN karma
- 221
- Public activity
- 39 items
- HN profile
- View on Hacker News ↗
About desideratum
No profile information was provided.
Recent public activity
-
comment
Comment #47658952
Yes my findings and thoughts were pretty much identical. I actually think you can get something reasonable at 1.3B params with the correct training recipe, but definitely not at th…
-
comment
Comment #47653535
This is a gross simplification of the process - you would typically use order(s) of magnitude more data and compute, and a substantial amount of online reinforcement learning to el…
-
comment
Comment #47653496
I appreciate the kind words very much : )
-
comment
Comment #47653490
I see what you mean, but I disagree. I expect that Claude Code is backed by a separate post-train of Claude base which has been trained using the Claude Code harness and toolset.
-
comment
Comment #47653455
Oh I wouldn't be surprised. This is a sample from one of the OSS code datasets I'd used, which are all generated synthetically using LLMs. Data is indeed the moat.
-
comment
Comment #47651050
This is a great question. You definitely aren't training this to use it, you're training it to understand how things work. It's an educational project, if you're interested in expe…
- story
- story
-
comment
Comment #47233320
Oh, and if you want to utilize 120Hz on the XDR display, you're going to have to replace your perfectly functioning Mac. > Mac models with M1, M1 Pro, M1 Max, M1 Ultra, M2, and M3 …
-
comment
Comment #47233169
It's mind-boggling that Apple is considering the base 27 inch Studio Display with the same 4 year old panel, but with some new accessories slapped on an "upgrade".
-
comment
Comment #46175566
Thanks for sharing this. I agree w.r.t. XLA. I've been moving to JAX after many years of using torch and XLA is kind of magic. I think torch.compile has quite a lot of catching up …
-
comment
Comment #46174520
The Scaling ML textbook also has an excellent section on TPUs. https://jax-ml.github.io/scaling-book/tpus/
-
comment
Comment #46088295
Aside: this guy regularly posts on the Discord server for an open-source post-training framework I maintain, demanding repayment for bugs in nightly builds and generally abusing th…
- story
- story
- story
- story
- story
- story
-
comment
Comment #42822495
This is an exceptional salary for the UK.
-
comment
Comment #42510728
I'd reccomend checking out the CUDA mode Discord server! They also have a channel for Metal https://discord.gg/ZqckTYcv
- story
-
comment
Comment #41690788
torchtune ( https://github.com/pytorch/torchtune ) - a PyTorch library for fine-tuning LLMs, particularly for memory-constrained setups. Try it out and fine-tune Llama3.1 8B on a s…
- story
- story