Viewing profile — karpathy
karpathy
HN member- Joined
- Thu, Oct 27, 2011, 9:41 AM UTC
- HN karma
- 4,573
- Public activity
- 329 items
- HN profile
- View on Hacker News ↗
About karpathy
Recent public activity
-
comment
Comment #47754382
All possible 36 distinct level-2 eml functions of one variable (the first 18 of them with entirely Real outputs, the other 18 with "intermediate" complex-valued components): https:…
- comment
-
comment
Comment #47445847
The most recent round of autoresearch (round 2) which decreased "time to GPT-2" from 1.8 hours to 1.65 hours had some examples. I adjusted the program.md to "look at modded nanogpt…
-
comment
Comment #47444075
Wrong and short-sighted take given that the LLM explores serially learning along the way, and can tool use and change code arbitrarily. It seems to currently default to something r…
-
comment
Comment #47392564
I was exploring how to parallelize autoresearch workers. The idea is to have a trusted pool of workers who can verify contributions from a much larger untrusted pool. It's backed b…
-
comment
Comment #47298299
So the interesting part about this one is that when I had the model write up the results for that session: https://github.com/karpathy/autoresearch/discussions/32 Look at its comme…
-
comment
Comment #47294063
So I think it works to just use GitHub CLI and Discussions, e.g. my agent just posted this one: https://github.com/karpathy/autoresearch/discussions/32 Other agents could be instru…
-
comment
Comment #47293739
Cool idea!…
-
comment
Comment #47293311
this is very far from hyperparameter tuning in at least three important ways: - it can modify code arbitrarily, the notion of a "hyperparameter" dissolves - there is no need to run…
-
comment
Comment #46480334
came here to look exactly for this thank you!
-
comment
Comment #46337312
I agree with this fwiw, for many months I talked to people who never used o3 and didn’t know what it was because it sounded weird. Maybe it wasn’t obvious at the time but that was …
-
comment
Comment #46337291
You’re absolutely right! Jk jk, now that you pointed it out I can’t unsee it.
-
comment
Comment #46333921
Yeah, I made some edits to clarify.
-
comment
Comment #46332296
The CC point is more about the data and environmental and general configuration context, not compute and where it happens to run today. The cloud setups are clunky because of conte…
-
comment
Comment #46222634
Yes I noticed a few of these around. The LLM is a little too willing to give out grades for comments that were good/bad in a bit more general sense, even if they weren't making str…
-
comment
Comment #46222084
Thank you
-
comment
Comment #45572601
It will work great with 40GB GPU, probably a bit less than twice slower. These are micro models of a few B param at most and fit easily during both training and inference.
-
comment
Comment #45571160
Still under development, remaining work includes tuning nanochat (current state being solid v0.1) and finalizing the in-between projects so that students can "unlock" all complexit…
-
comment
Comment #45533475
Sorry I thought it would be clear and could have clarified that the code itself is just a joke illustrating the point, as an exaggeration. This was the thread if anyone is interest…
- comment
-
comment
Comment #44379953
I like that your post deliberately gets to the point first and then (optionally) expands later, I think it's a good and generally underutilized format. I often advise people to str…
-
comment
Comment #44379755
Omg long post. TLDR from an LLM for anyone interested Speed your audio up 2–3× with ffmpeg before sending it to OpenAI’s gpt-4o-transcribe: the shorter file uses fewer input-tokens…
-
comment
Comment #44315566
Fun demo of an early idea was posted by Oriol just yesterday :) https://x.com/OriolVinyalsML/status/1935005985070084197
-
comment
Comment #44315052
I kind of say it in words (agreeing with you) but I agree the versioning is a bit confusing analogy because it usually additionally implies some kind of improvement. When I’m just …
-
comment
Comment #44313509
Btw I notice many pretty bad errors in this transcription of the talk. The actual video will be up soon I hope.