Live data from Hacker News

FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

epochai.org

91–100 of 110 posts

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#91

Earlier quoted context omitted.

An AGI is often claimed to be a general purpose problem solver and these are exactly the types of problems that a general purpose problem solver would be able to solve if given access to a mathematical library. All existing LLMs have been trained on abstract mathematics and logic but it is obvious that they are incapable of abstract logical reasoning, e.g. solving sudoku puzzles.

Here is my prediction, FWIW: the hard part of the problem has already been solved, in the following technical sense: there is a few 1000 lines program that has not been invented yet, but it will be invented soon, that loads a current LLM model, runs fast on current hardware, and you will deem it to be an AGI. In other words, the conditional Kolmogorov complexity of undisputable AGI given the Llama weights is only a f…

[dead]

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#92

Earlier quoted context omitted.

An AGI is often claimed to be a general purpose problem solver and these are exactly the types of problems that a general purpose problem solver would be able to solve if given access to a mathematical library. All existing LLMs have been trained on abstract mathematics and logic but it is obvious that they are incapable of abstract logical reasoning, e.g. solving sudoku puzzles.

Here is my prediction, FWIW: the hard part of the problem has already been solved, in the following technical sense: there is a few 1000 lines program that has not been invented yet, but it will be invented soon, that loads a current LLM model, runs fast on current hardware, and you will deem it to be an AGI. In other words, the conditional Kolmogorov complexity of undisputable AGI given the Llama weights is only a f…

I think you are likely right but coming up with that final inference strategy is still "the hard part" IMO. Not in terms of computation, but in terms of algorith development.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#93
post #57

Earlier quoted context omitted.

I think because if you end up having an AI that is as capable as the graduate students Tao is used to dealing with (so basically potential field medalists) then you are basically betting that 85% chance something like AGI (at least in consequence) will be here in 3 years. It is possible, but 85% chance?

>then you are basically betting that 85% chance something like AGI Not really. It would just need to do more steps in a sequence that current models do. And that number has been going up consistently. So it would be just another narrow AI expert system. It is very likely that it will be solved, but it is very unlikely that it will be generally capable in the sense most researchers understand AGI today.

I am willing to bet it won't be solved by 2028 and the betting market is overestimating AI capabilities and progress on abstract reasoning. No current AI on the market can consistently synthesize code according to a logical specification and that is almost certainly a requirement for solving this benchmark.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#94
post #51

Earlier quoted context omitted.

Simple: I highly doubt they're willing to risk a scandal that would further tarnish their brand. It's still reeling from last year's drama, in addition to a spate of high-profile departures this year. Not to mention a few articles with insider sources that aren't exactly flattering.

You’re not thinking about the other side of the equation. If they win (becoming the first to excel at the benchmark), they potentially make billions. If they lose, they’ll be relegated to the dustbin of LLM history. Since there is an existential threat to the brand, there is almost nothing that isn’t worth risking to win. Risking a scandal to avoid irrelevance is an easy asymmetrical bet. Of course they would take th…

Okay, let's assume what you say ends up being true. They effectively cheat, then raise some large fundraising round predicated on those results.

Two months later there's a bombshell exposé detailing insider reports of how they cheated the test by cooking their training data using an army of PhDs to hand-solve. Shame.

At a minimum investor confidence goes down the drain, if it doesn't trigger lawsuits from their investors. Then you're looking at maybe another CEO ouster fiasco with a crisis of faith across their workforce. That workforce might be loyal now, but that's because their RSUs are worth something and not tainted by fraud allegations.

If you're right, I suppose it really depends on how well they could hide it via layers of indirection and compartmentalization, and how hard they could spin it. I don't really have high hopes for that given the number of folks there talking to the press lately.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#95

Earlier quoted context omitted.

Since you go that route, do you have published evidence that shows they HAVENT entered the top of the S-curve?

For one, thinking LLMs have plateaued is essentially assuming that video can't teach AI anything. It's like saying a person locked into a room his whole life with only books to read would be as good at reasoning as someone's who's been out in the world.

LLMs do not learn the same way that a person does

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#96

Earlier quoted context omitted.

I have yet to see any published evidence of that.

Since you go that route, do you have published evidence that shows they HAVENT entered the top of the S-curve?

This is not a linear process. Deep-learning models do not scale that way.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#97
post #27

Earlier quoted context omitted.

Sure but it is also reasonable to consider that the pace of progress is not always exponential or even linear at best. Diminishing returns are a thing and we already know that a 405b model is not 5 times better than a 70b model.

How do you determine the multiplier. Because e.g. there are many problems that GPT4 can solve while GPT3.5 can't. In this case it is infinitely better.

Let's say your benchmark gets you at 60% with a 70b parameter model and you get to 65% with a 405b one, it's fairly obvious that it's just incremental progress, not a sustainable growth of capabilities per added parameter. Also, most of the data used these days for trainings these very large models is synthetic data, which is probably very low quality overall compared to human-sourced data.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#98

Not very impressed by the problems they displayed but I guess there should be some good problems in the set given the comments (not in the sense that I find them super easy but they seems random and not super well-posed, and extremely artificial problems--in the sense that they seem to not be of particular mathematical interest[or at least the mathematical content of the problem is being deliberately hidden for testi…

I guess the primary reason is that the answers must be numbers that can be verified easily. Otherwise, you just flood the validator with long LLM reasoning that's hard to verify. People have been proposing using LEAN as a medium for answers but AFAIK even LEAN is not mainstream in the general math community, so there's always trade-offs.

Also, coming up with good problems is an art in its own right; the Soviets was famous for institutionalizing anti-Semitism via special math puzzles for Jews in Moscow Univerisity entrance exams. The questions are constructed as such that are hard to solve but have some elementary solutions to divert criticism.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#99
post #17

Earlier quoted context omitted.

It's my job to build these tools. I'm well aware of their strengths and shortcomings.

Unless you are building one of the frontier models, I’m not sure that your experience gives you insight on those models. Perhaps it just creates needless assumptions.

I'm building the layer on top of the models.

People call it agentic AI but that's a word without a definition.

Needless to say better LLMs help with my work, the same way that a stronger horse makes plowing easier.

The difference is that unlike horse breeders, e.g. Anthropic and openAI, I want to get to the internal combustion engine and tractors.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#100

Earlier quoted context omitted.

>That's a definition some people use yes but a machine that can solve any problem any human can is by definition super-intelligent and super-capable because there exists no human that can solve any problem any human can. We don't need every human in the world to learn complex topology math like Terence Tao. Some need to be farmers. Some need to be engineers. Some need to be kindergarten teachers. When we need someone…

>We don't need every human in the world to learn complex topology math like Terence Tao. Some need to be farmers. Some need to be engineers. Some need to be kindergarten teachers. It doesn't have much to do with need. Not every human can be as capable regardless of how much need or time you allocate for them to do so. Then some humans are shoulders above peers in one field but come a bit short in another closely rela…

There is no one true definition but your definition is definitely the less popular one.
Post reply on HN