Live data from Hacker News

I trained a small transformer in 1.5hrs and it beats many LLMs

mvakde.github.io

61–70 of 183 posts

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#61
post #44

Earlier quoted context omitted.

>> NOT an LLM. its a small ar transforme Super cool project! Though, aren't most modern LLM's ar transformers internally?

It's an slm

There is no language in the training of this, so there is no l.

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#62

Hi! Author here. Surprised to see this on HN now. Happy to answer any questions! Some context about this: - This is NOT an LLM. its a small ar transformer trained from scratch. One of the points was that extremely complex problems can be tackled without LLMs - Till the v1 of this result, this benchmark was only scaled by LLMs or their finetunes (ofc w enormous training costs). Other attempts performed okayish but use…

First: this is really great technical writing, especially when you get into the rebuttals. Firm & clear without polemics -- props, and thanks for open-sourcing!

That said; I don't have the time, energy, or anywhere near the expertise to challenge you on the DL specifics, but I feel compelled to add another voice to the chorus of doubters nonetheless. Using other ARC examples at runtime (effectively, yes?) for "transduction" may not violate what some officer behind ARC said on Twitter --and is certainly a fantastic tool for certain problem spaces-- but it just seems like a glaring and unavoidable philosophical problem in this one. My issue isn't with using the eval set per-se (though that obviously sets off well-justified alarm bells), but rather building an AGI system whose performance relies on the arbitrary size and shape of this particular dataset.

There's a lot of ways to frame this, but given the transduction citations the most appropriate is probably the AI winter's infamous 'Frame Problem':

You say upfront that this works in the first place because ARC has "very few samples... in a high dimensional space"; to me, that seems like an extremely strong indicator that the datasets are not intended to capture anywhere near the full semantic space that we would consider relevant for AGI. If true, your approach would indeed be ""cheating"" by using an arbitrary & unavoidable feature of the dataset (that they didn't have the time or money to craft 100 million high quality cases by hand instead of 1000) to solve the frame problem upfront for you. This would explain why you don't even need a full LLM here -- that's the unsolvable problem that LLMs solve for us.

That is... even if the ARC train+eval sets contain the sum of human intuition between them, superficial differences IRL would render your model unable to identify which examples are relevant to which problems, and thus unable to transductively reason.

In plainer English: surely you'd agree that your model would do worse if we swapped it out with Opus behind the scenes than the next-highest-scoring ARC model would do in the same position, yes? For coding, research, dumb questions, SVG pelicans -- the lot?

If so, that seems like hard proof that this scores high on a benchmark at the cost of the benchmark itself. Like, if this transductive approach leads to ARC1 being claimed (which I thought it was ages ago but :shrug:), they'll either have to abandon the whole benchmark or ban this approach retroactively.

If not... well, I guess I encourage you to try it! It seems like you'd need 1000 truly stellar hand-picked examples to transductively cover that whole space, for one thing.

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#63
post #55

I think(?) you’ve already probably done a good job of explaining this criticism for semi-informed people. But can you dumb it down even more for those of us who are almost entirely out-of-the-loop? > Training on the eval puzzles is cheating / “training on test” > No this is false. “Training on test” specifically means training on the labels of test data. The labels were not trained on. > Also, ARC is a metalearning b…

The point of ARC is essentially an "IQ Test" for AI systems. It is meant to cover abstract reasoning capabilities of generally-intelligent systems like LLMs. What the author did here was build a system that only solves ARC problems. The other tension is the fact that this score is on the public eval set. In machine learning, you typically have 3 datasets: training, evaluation, and test. The training set is the datase…

It seems like an interesting strategy. Based on the author’s comment, they haven’t been at it for very long. So, I guess the folks who run the private test haven’t had a chance to get to it? It’d be interesting to hear how it does.

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#64
post #50
post #44

Earlier quoted context omitted.

>> NOT an LLM. its a small ar transforme Super cool project! Though, aren't most modern LLM's ar transformers internally?

Not all transformers are _language_ models - the sequences of tokens don't have to be sequences of words.

In this case, what are the tokens?

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#65
post #6

Is the author only running their model against one benchmark? I don't think anyone finds that difficult to achieve, the difficulty comes when you want to make the model not benchmaxxed to a specific benchmark, and generalize so it can solve problems not part of the training data, but seems this model is specifically for not this? How useful is that? If you just wanted to pass these specific tasks in this specific ben…

I read this as a response to the current hype around LLMs. He is showing computers can solve these issues, without using an LLM architecture. A lot of people have sort of forgot that machine learning is more than just LLMs these days. I found it to be a very interesting angle.

Just to be clear: It was well known that you can reach such scores with small models and without an LLM if you train on the task. The author highlights those models himself - e.g. HRM/TRM.

The novelty is more that it works with such a plain transformer and low compute price.

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#66

I think(?) you’ve already probably done a good job of explaining this criticism for semi-informed people. But can you dumb it down even more for those of us who are almost entirely out-of-the-loop? > Training on the eval puzzles is cheating / “training on test” > No this is false. “Training on test” specifically means training on the labels of test data. The labels were not trained on. > Also, ARC is a metalearning b…

What you gather is correct, assuming by "the test" you mean the ARC benchmark in general. It was controversial because people are used to LLMs which are frozen at train time, where the eval problems are usually not trained on for various reasons like fragility (basically porridgeraisin's ans which is great)

Here's another explanation. Take the train dataset and test dataset of a benchmark

Train: {x_i -> f(x_i)}, Test: {x_j -> f(x_j)}

As long as f(x_j) in the test set is hidden, there is no "training on test". In a normal benchmark, each x_i is a single datapoint. But in metalearning benchmarks like ARC, x_i is the puzzle itself that has a train set and the test questions within it, hence the confusion and controversy

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#67
post #11

Earlier quoted context omitted.

Crazy, considering rhabdo isn't that rare.

This was in India, where he describes the medical knowledge of providers as subpar at best.

Yeah but rhabdo is something literally any e.g. body builder, power lifter, etc could tell you about. Actually if somebody knows what hypertrophy is, they probably know what rhabdo is. It's a pretty normal and big concern in any sort of high intensity weight training.

I can't think of many ways that otherwise healthy and fit younger people can physically nearly kill themselves doing normal activity, so it kind of stands out - let alone it being not all that rare either. Rhabdo has even gone viral in the news like when a while back a couple of Chinese girls nearly killed themselves doing a social media 'squat challenge.' They did 1000, got rhabdo, didn't know what was happening, ended up in the ICU with kidney damage.

It also manifests in other ways too. For instance I had an elderly family member give himself rhabdo during a manic phase he was going through when he started going wild on construction and other physical tasks that were way beyond what his body was ready for.

Basically it's not some super obscure thing you'd expect only a good doctor, let alone a specialist, to know about.

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#68
post #55

Earlier quoted context omitted.

The point of ARC is essentially an "IQ Test" for AI systems. It is meant to cover abstract reasoning capabilities of generally-intelligent systems like LLMs. What the author did here was build a system that only solves ARC problems. The other tension is the fact that this score is on the public eval set. In machine learning, you typically have 3 datasets: training, evaluation, and test. The training set is the datase…

It seems like an interesting strategy. Based on the author’s comment, they haven’t been at it for very long. So, I guess the folks who run the private test haven’t had a chance to get to it? It’d be interesting to hear how it does.

I'm currently 10th in the world on the private set on Kaggle. And iirc, at one point I was 4th

Can't comment more since its an ongoing competition

Re: I trained a small transformer in 1.5hrs and it beats many LLMs

#69

I think(?) you’ve already probably done a good job of explaining this criticism for semi-informed people. But can you dumb it down even more for those of us who are almost entirely out-of-the-loop? > Training on the eval puzzles is cheating / “training on test” > No this is false. “Training on test” specifically means training on the labels of test data. The labels were not trained on. > Also, ARC is a metalearning b…

Basically, you have a bunch of Q,A pairs in the training dataset. Here, it was trained to next-word predict the question itself, as well as next-word predict the answer given the question as prompt. This is bog-standard, no one's complaining. In the test dataset's Q,A pairs, it was only trained to next-word predict the question itself, and it was not given the answer at all. It was then evaluated by seeing if it is a…

Further question—the model produces an answer to the question, it sends the answer, and then gets graded. Does it get to know immediately how it did, or does it get the grade back at the end after answering all the questions?

If it is the former case, it would be possible to add the generated question/answer pair into the training set as well. Would that be considered fair? (Of course this is a moot point if the answers all get graded simultaneously at the end). Then the model could explore interesting strategies around what order to answer questions in.

In my uninformed opinion, the various permutations of question ordering/answer revealing all map to different real-world scenarios… and any of them could be interesting!

Post reply on HN