Live data from Hacker News

FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

epochai.org

71–80 of 110 posts

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#71
post #6
post #4

Earlier quoted context omitted.

> they’re merely regurgitating memorized information Source?

I'm not sure if it is feasible to provide all relevant sources to someone who doesn't follow a field. It is quite common knowledge that LLMs in their current form have no ability to recurse directly over a prompt, which inherently limits their reasoning ability.

I am not looking for all sources. And I do follow the field. I just don’t know the sources that would back the claim they are making. Nor do I understand why limits on recursion means there is no reasoning and only memorization.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#72
post #17

Earlier quoted context omitted.

It’s sometimes like, are these critics using the tools? It’s a strange schism at the moment.

It's my job to build these tools. I'm well aware of their strengths and shortcomings.

Unless you are building one of the frontier models, I’m not sure that your experience gives you insight on those models. Perhaps it just creates needless assumptions.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#73
post #57

Earlier quoted context omitted.

> Surprisingly, prediction markets [1] are putting 62% on AI achieving > 85% performance on the benchmark before 2028. Why surprisingly? 2028 is twice as long as capable LLMs existed to date. By "capable" here I mean capable enough to even remotely consider the idea of LLMs solving such tasks in the first place. ChatGPT/GPT-3.5 isn't even 2 years old ! 4 years is a lot of time. It's kind of silly to assume LLM capabi…

I think because if you end up having an AI that is as capable as the graduate students Tao is used to dealing with (so basically potential field medalists) then you are basically betting that 85% chance something like AGI (at least in consequence) will be here in 3 years. It is possible, but 85% chance?

It would also require ability to easily handle large amount of complex information and dependencies such as massive codebases etc and then also be able to operate physically like humans do. By controlling a robot of some sort.

Being able to solve self contained exercise can be obviously very challenging, but there are other different types of skills that might or might not be related and have to be solved as well.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#74

Earlier quoted context omitted.

I have yet to see any published evidence of that.

Since you go that route, do you have published evidence that shows they HAVENT entered the top of the S-curve?

Why are you thinking in binary. It is not clear at all to me that the progress is stagnating, and in fact I am still impressed by the progress. But I couldn't tell whether there is going to come a wall or not. There is no clear reason why there should be some sort of standard or historical curve for this progress.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#75
post #30

Earlier quoted context omitted.

If I was going to bet, I would bet yes, they will reach above 85% performance. The problem with all benchmarks, one that we just don't how to solve, is leakage. Systematically, LLMs are much better at benchmarks created before they were trained than after. There are countless papers that show significant leakage between training and test sets for models. This is in part why so many LLMs are so strong according to ben…

This benchmark’s questions and answers will be kept fully private, and the benchmark will only be run by Epoch. Short of the companies fishing out the questions from API logs (which seems quite unlikely), this shouldn’t be a problem.

Ideally they would have batches of those exercises, where the only use the next batch when someone has solved a suspicious amount of those exercises. If it performs much worse on the next batch, that is a tell of leakage.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#76
post #27

Earlier quoted context omitted.

> Surprisingly, prediction markets [1] are putting 62% on AI achieving > 85% performance on the benchmark before 2028. Why surprisingly? 2028 is twice as long as capable LLMs existed to date. By "capable" here I mean capable enough to even remotely consider the idea of LLMs solving such tasks in the first place. ChatGPT/GPT-3.5 isn't even 2 years old ! 4 years is a lot of time. It's kind of silly to assume LLM capabi…

Sure but it is also reasonable to consider that the pace of progress is not always exponential or even linear at best. Diminishing returns are a thing and we already know that a 405b model is not 5 times better than a 70b model.

How do you determine the multiplier. Because e.g. there are many problems that GPT4 can solve while GPT3.5 can't. In this case it is infinitely better.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#77
post #18

Earlier quoted context omitted.

It will be a useful benchmark to validate claims by people like Sam Altman about having achieved AGI.

Most humans can't solve these problems, so it's certainly possible to imagine a legitimate AGI that can't either.

AGI should be able to do anything the best humans can do. ASI is when it does everything better than the best humans.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#78
post #62

Earlier quoted context omitted.

I think people just see FrontierMath as a goal post that an AGI needs to hit. The term "artificial general intelligence" implies that it can solve any problem a human can. If it can't solve math problems that an expert human can, then it's not AGI by definition. I think we have to keep in mind that humans have specialized. Some do law. Some do math. Some are experts at farming. Some are experts at dance history. It's…

Okay, sounds like different definitions. If you have a single system that can solve any problem any human can, I'd call that ASI, as it's way smarter than any human. It's an extremely high bar, and before we reach it I think we'll have very intelligent systems that can do more than most humans, so it seems strange not to call those AGIs (they would meet the definition of AGI on Wikipedia [1]). [1] https://en.wikipedi…

The reason for the AGI definition is to indicate a point where no human can provide more value than the AGI can. AGI should be able to replace all work efforts on its own, as long as it can scale.

ASI is when it is able to develop a much better version of itself to then iteratively go past all of that.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#79
post #62

Earlier quoted context omitted.

Okay, sounds like different definitions. If you have a single system that can solve any problem any human can, I'd call that ASI, as it's way smarter than any human. It's an extremely high bar, and before we reach it I think we'll have very intelligent systems that can do more than most humans, so it seems strange not to call those AGIs (they would meet the definition of AGI on Wikipedia [1]). [1] https://en.wikipedi…

>If you have a single system that can solve any problem any human can, I'd call that ASI I don't think that's the popular definition. AGI = solve any problem any human can. In this case, we've not reached AGI since it can't solve most FrontierMath problems. ASI = intelligence far surpasses even the smartest humans. If the definition of AGI has is that it's more intelligent than the average human, you can argue that w…

>I don't think that's the popular definition.

People have all sorts of definitions for AGI. Some are more popular than others but at this point, there is no one true definition. Even Open AI's definition is different from what you have just said. They define it as "highly autonomous systems that outperform humans in most economically valuable tasks"

>AGI = solve any problem any human can.

That's a definition some people use yes but a machine that can solve any problem any human can is by definition super-intelligent and super-capable because there exists no human that can solve any problem any human can.

>If the definition of AGI has is that it's more intelligent than the average human, you can argue that we already have AGI today. But no one thinks we have AGI today.

There are certainly people who do, some of which are pretty well respected in the community, like Norvig.

https://www.noemamag.com/artificial-general-intelligence-is-...

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#80

Earlier quoted context omitted.

I mean, this benchmark is really hard. I don't think it's a requirement that a system claiming to be AGI should be able to solve these problems, 99.99% of humans can't either.

An AGI is often claimed to be a general purpose problem solver and these are exactly the types of problems that a general purpose problem solver would be able to solve if given access to a mathematical library. All existing LLMs have been trained on abstract mathematics and logic but it is obvious that they are incapable of abstract logical reasoning, e.g. solving sudoku puzzles.

99.9% of the population would not be able to solve these problems given a year and access to every piece of mathematical literature ever written (except the solutions to these problems of course).

Saying that you need to solve these to be considered AGI is ridiculously strict.

Post reply on HN