Live data from Hacker News

Viewing profile — desideratum

desideratum

HN member
Joined
Wed, Jul 24, 2019, 9:24 PM UTC
HN karma
221
Public activity
39 items

About desideratum

No profile information was provided.

Recent public activity

  1. comment
    Comment #47658952

    Yes my findings and thoughts were pretty much identical. I actually think you can get something reasonable at 1.3B params with the correct training recipe, but definitely not at th…

  2. comment
    Comment #47653535

    This is a gross simplification of the process - you would typically use order(s) of magnitude more data and compute, and a substantial amount of online reinforcement learning to el…

  3. comment
    Comment #47653496

    I appreciate the kind words very much : )

  4. comment
    Comment #47653490

    I see what you mean, but I disagree. I expect that Claude Code is backed by a separate post-train of Claude base which has been trained using the Claude Code harness and toolset.

  5. comment
    Comment #47653455

    Oh I wouldn't be surprised. This is a sample from one of the OSS code datasets I'd used, which are all generated synthetically using LLMs. Data is indeed the moat.

  6. comment
    Comment #47651050

    This is a great question. You definitely aren't training this to use it, you're training it to understand how things work. It's an educational project, if you're interested in expe…

  7. story
  8. story
  9. comment
    Comment #47233320

    Oh, and if you want to utilize 120Hz on the XDR display, you're going to have to replace your perfectly functioning Mac. > Mac models with M1, M1 Pro, M1 Max, M1 Ultra, M2, and M3 …

  10. comment
    Comment #47233169

    It's mind-boggling that Apple is considering the base 27 inch Studio Display with the same 4 year old panel, but with some new accessories slapped on an "upgrade".

  11. comment
    Comment #46175566

    Thanks for sharing this. I agree w.r.t. XLA. I've been moving to JAX after many years of using torch and XLA is kind of magic. I think torch.compile has quite a lot of catching up …

  12. comment
    Comment #46174520

    The Scaling ML textbook also has an excellent section on TPUs. https://jax-ml.github.io/scaling-book/tpus/

  13. comment
    Comment #46088295

    Aside: this guy regularly posts on the Discord server for an open-source post-training framework I maintain, demanding repayment for bugs in nightly builds and generally abusing th…

  14. story
  15. story
  16. story
  17. story
  18. story
  19. story
  20. comment
    Comment #42822495

    This is an exceptional salary for the UK.

  21. comment
    Comment #42510728

    I'd reccomend checking out the CUDA mode Discord server! They also have a channel for Metal https://discord.gg/ZqckTYcv

  22. story
  23. comment
    Comment #41690788

    torchtune ( https://github.com/pytorch/torchtune ) - a PyTorch library for fine-tuning LLMs, particularly for memory-constrained setups. Try it out and fine-tune Llama3.1 8B on a s…

  24. story
  25. story