I trained a small transformer in 1.5hrs and it beats many LLMs
121–130 of 183 posts
Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#122Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#123Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#124Earlier quoted context omitted.
Yeah but rhabdo is something literally any e.g. body builder, power lifter, etc could tell you about. Actually if somebody knows what hypertrophy is, they probably know what rhabdo is. It's a pretty normal and big concern in any sort of high intensity weight training. I can't think of many ways that otherwise healthy and fit younger people can physically nearly kill themselves doing normal activity, so it kind of sta…
People commonly knowing about Rhabdo is a much newer thing. I never heard people talk regularly about it all before CrossFit become popular
Still, emergency doctors, especially in a hot place like India, should be aware of it.
Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#125Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#126Hi! Author here. Surprised to see this on HN now. Happy to answer any questions! Some context about this: - This is NOT an LLM. its a small ar transformer trained from scratch. One of the points was that extremely complex problems can be tackled without LLMs - Till the v1 of this result, this benchmark was only scaled by LLMs or their finetunes (ofc w enormous training costs). Other attempts performed okayish but use…
Thank you for this excellent post series. It reminds me a lot of the pre-LLM days, though I was mostly using LSTMs back then. When the original GPT paper came out, I thought the future would be using LLMs to generate tons of synthetic labeled data and then training specialized LSTM or transformer models per-task. Had a couple of questions: 1) You note that ARC-AGI is a meta-learning task, have you tried any meta-lear…
1) Unfortunately I didn't. I was v new to ML when I did this and didnt have time or skill to try many things. Will try them when I get some time!
2) Possibly, but it would require significant changes and effort. But much larger models would be required imo (must have capacity greater than the complexity of the problem)
Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#127Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#128Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#129> Ban offline training/pretraining. Models must train from scratch after submission Previously this was considered impossible so rule. My model shows this is possible Guarantees no synthetic data can be used It makes the comparison fair across differet models. Otherwise some models like LLMs can benchmaxx ARC by using ungodly amounts of offline training. (Since the benchmark has been around a long time, many ARC-like…
Other competitions have implemented things like this before. Eg: OpenAI's Parameter Golf and Keller Jordan's Modded NanoGPT Speedrun
Re: I trained a small transformer in 1.5hrs and it beats many LLMs
#130Earlier quoted context omitted.
I studied biochem in undergrad and my classes were full of premed students. I loved the subject and nerded out about the course material - I spent my time designing my own experiments around gene cloning that took several semesters to run. They were sharing last year's tests with their frat buddies and laughing at us nerds. I've never looked at doctors the same way again after college. I looked up to them as a child,…
>When we should let doctors from overseas immigrate and easily become practicing doctors here in the US. Lol, what does this has to do with anything regarding aptitude or curiosity!?