Live data from Hacker News

Viewing profile — karpathy

karpathy

HN member
Joined
Thu, Oct 27, 2011, 9:41 AM UTC
HN karma
4,573
Public activity
329 items

About karpathy

Independent researcher. Previously Director of AI at Tesla, OpenAI, CS231n, PhD @ Stanford. I like to train large deep neural nets

Recent public activity

  1. comment
    Comment #47754382

    All possible 36 distinct level-2 eml functions of one variable (the first 18 of them with entirely Real outputs, the other 18 with "intermediate" complex-valued components): https:…

  2. comment
  3. comment
    Comment #47445847

    The most recent round of autoresearch (round 2) which decreased "time to GPT-2" from 1.8 hours to 1.65 hours had some examples. I adjusted the program.md to "look at modded nanogpt…

  4. comment
    Comment #47444075

    Wrong and short-sighted take given that the LLM explores serially learning along the way, and can tool use and change code arbitrarily. It seems to currently default to something r…

  5. comment
    Comment #47392564

    I was exploring how to parallelize autoresearch workers. The idea is to have a trusted pool of workers who can verify contributions from a much larger untrusted pool. It's backed b…

  6. comment
    Comment #47298299

    So the interesting part about this one is that when I had the model write up the results for that session: https://github.com/karpathy/autoresearch/discussions/32 Look at its comme…

  7. comment
    Comment #47294063

    So I think it works to just use GitHub CLI and Discussions, e.g. my agent just posted this one: https://github.com/karpathy/autoresearch/discussions/32 Other agents could be instru…

  8. comment
    Comment #47293739

    Cool idea!…

  9. comment
    Comment #47293311

    this is very far from hyperparameter tuning in at least three important ways: - it can modify code arbitrarily, the notion of a "hyperparameter" dissolves - there is no need to run…

  10. comment
    Comment #46480334

    came here to look exactly for this thank you!

  11. comment
    Comment #46337312

    I agree with this fwiw, for many months I talked to people who never used o3 and didn’t know what it was because it sounded weird. Maybe it wasn’t obvious at the time but that was …

  12. comment
    Comment #46337291

    You’re absolutely right! Jk jk, now that you pointed it out I can’t unsee it.

  13. comment
    Comment #46333921

    Yeah, I made some edits to clarify.

  14. comment
    Comment #46332296

    The CC point is more about the data and environmental and general configuration context, not compute and where it happens to run today. The cloud setups are clunky because of conte…

  15. comment
    Comment #46222634

    Yes I noticed a few of these around. The LLM is a little too willing to give out grades for comments that were good/bad in a bit more general sense, even if they weren't making str…

  16. comment
    Comment #46222084

    Thank you

  17. comment
    Comment #45572601

    It will work great with 40GB GPU, probably a bit less than twice slower. These are micro models of a few B param at most and fit easily during both training and inference.

  18. comment
    Comment #45571160

    Still under development, remaining work includes tuning nanochat (current state being solid v0.1) and finalizing the in-between projects so that students can "unlock" all complexit…

  19. comment
    Comment #45533475

    Sorry I thought it would be clear and could have clarified that the code itself is just a joke illustrating the point, as an exaggeration. This was the thread if anyone is interest…

  20. comment
  21. comment
    Comment #44379953

    I like that your post deliberately gets to the point first and then (optionally) expands later, I think it's a good and generally underutilized format. I often advise people to str…

  22. comment
    Comment #44379755

    Omg long post. TLDR from an LLM for anyone interested Speed your audio up 2–3× with ffmpeg before sending it to OpenAI’s gpt-4o-transcribe: the shorter file uses fewer input-tokens…

  23. comment
    Comment #44315566

    Fun demo of an early idea was posted by Oriol just yesterday :) https://x.com/OriolVinyalsML/status/1935005985070084197

  24. comment
    Comment #44315052

    I kind of say it in words (agreeing with you) but I agree the versioning is a bit confusing analogy because it usually additionally implies some kind of improvement. When I’m just …

  25. comment
    Comment #44313509

    Btw I notice many pretty bad errors in this transcription of the talk. The actual video will be up soon I hope.