Live data from Hacker News

FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

epochai.org

81–90 of 110 posts

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#81

Earlier quoted context omitted.

>If you have a single system that can solve any problem any human can, I'd call that ASI I don't think that's the popular definition. AGI = solve any problem any human can. In this case, we've not reached AGI since it can't solve most FrontierMath problems. ASI = intelligence far surpasses even the smartest humans. If the definition of AGI has is that it's more intelligent than the average human, you can argue that w…

>I don't think that's the popular definition. People have all sorts of definitions for AGI. Some are more popular than others but at this point, there is no one true definition. Even Open AI's definition is different from what you have just said. They define it as "highly autonomous systems that outperform humans in most economically valuable tasks" >AGI = solve any problem any human can. That's a definition some peo…

>That's a definition some people use yes but a machine that can solve any problem any human can is by definition super-intelligent and super-capable because there exists no human that can solve any problem any human can.

We don't need every human in the world to learn complex topology math like Terence Tao. Some need to be farmers. Some need to be engineers. Some need to be kindergarten teachers. When we need someone to solve those problems, we can call Terence Tao.

When AI needs to solve those problems, it can't do it without humans in 2024. Period.

That's the whole point of this discussion.

The definition of ASI historically is that it's an intelligence that far surpasses humans - not at the level of the best humans.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#82
post #51
post #39

Earlier quoted context omitted.

Why do people still insist that this is unlikely? Like assuming that the company that payed 15M for chat.com does not have some spare change to pay some graduate students/postdocs to solve some math problems. The publicity of solving such benchmark would definitely raise the valuation so it would 100% be worth it for them...

Simple: I highly doubt they're willing to risk a scandal that would further tarnish their brand. It's still reeling from last year's drama, in addition to a spate of high-profile departures this year. Not to mention a few articles with insider sources that aren't exactly flattering.

Parallel construction

Doesnt cause too much scandal lol

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#83

Earlier quoted context omitted.

>I don't think that's the popular definition. People have all sorts of definitions for AGI. Some are more popular than others but at this point, there is no one true definition. Even Open AI's definition is different from what you have just said. They define it as "highly autonomous systems that outperform humans in most economically valuable tasks" >AGI = solve any problem any human can. That's a definition some peo…

>That's a definition some people use yes but a machine that can solve any problem any human can is by definition super-intelligent and super-capable because there exists no human that can solve any problem any human can. We don't need every human in the world to learn complex topology math like Terence Tao. Some need to be farmers. Some need to be engineers. Some need to be kindergarten teachers. When we need someone…

>We don't need every human in the world to learn complex topology math like Terence Tao. Some need to be farmers. Some need to be engineers. Some need to be kindergarten teachers.

It doesn't have much to do with need. Not every human can be as capable regardless of how much need or time you allocate for them to do so. Then some humans are shoulders above peers in one field but come a bit short in another closely related one they've sunk a lot of time into.

Like i said, arguing about a one true definition is pointless. It doesn't exist.

>The definition of ASI historically is that it's an intelligence that far surpasses humans - not at the level of the best humans.

A Machine that is expert level in every single field would likely far surpass the output of any human very quickly. Yes, there might exist intelligences that are significantly more 'super' but that is irrelevant. Competence, like generality is a spectrum. You can have two super-human intelligences with a competence gap.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#84
post #60

Earlier quoted context omitted.

> Surprisingly, prediction markets [1] are putting 62% on AI achieving > 85% performance on the benchmark before 2028. Why surprisingly? 2028 is twice as long as capable LLMs existed to date. By "capable" here I mean capable enough to even remotely consider the idea of LLMs solving such tasks in the first place. ChatGPT/GPT-3.5 isn't even 2 years old ! 4 years is a lot of time. It's kind of silly to assume LLM capabi…

People really love pointing at the first part of a logistic curve and go "behold! an exponential".

I think (to give them the most generous read) they are just betting the halfway is still pretty far ahead. It is a different bet but IMO not an inherently ridiculous one like just misidentifying the shape of the thing; everything is a logistic curve, right? At least, everything that doesn’t blow up to infinity.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#85
post #57

Earlier quoted context omitted.

> Surprisingly, prediction markets [1] are putting 62% on AI achieving > 85% performance on the benchmark before 2028. Why surprisingly? 2028 is twice as long as capable LLMs existed to date. By "capable" here I mean capable enough to even remotely consider the idea of LLMs solving such tasks in the first place. ChatGPT/GPT-3.5 isn't even 2 years old ! 4 years is a lot of time. It's kind of silly to assume LLM capabi…

I think because if you end up having an AI that is as capable as the graduate students Tao is used to dealing with (so basically potential field medalists) then you are basically betting that 85% chance something like AGI (at least in consequence) will be here in 3 years. It is possible, but 85% chance?

>then you are basically betting that 85% chance something like AGI

Not really. It would just need to do more steps in a sequence that current models do. And that number has been going up consistently. So it would be just another narrow AI expert system. It is very likely that it will be solved, but it is very unlikely that it will be generally capable in the sense most researchers understand AGI today.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#86
post #30

Earlier quoted context omitted.

If I was going to bet, I would bet yes, they will reach above 85% performance. The problem with all benchmarks, one that we just don't how to solve, is leakage. Systematically, LLMs are much better at benchmarks created before they were trained than after. There are countless papers that show significant leakage between training and test sets for models. This is in part why so many LLMs are so strong according to ben…

This benchmark’s questions and answers will be kept fully private, and the benchmark will only be run by Epoch. Short of the companies fishing out the questions from API logs (which seems quite unlikely), this shouldn’t be a problem.

> answers will be kept fully private

> Short of the companies fishing out the questions from API logs (which seems quite unlikely)

They all pretty clearly state[1] versions of "We use your queries (removing personal data) to improve the models" so I'm not sure why that's unlikely.

https://help.openai.com/en/articles/5722486-how-your-data-is...

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#87
post #51
post #39

Earlier quoted context omitted.

Why do people still insist that this is unlikely? Like assuming that the company that payed 15M for chat.com does not have some spare change to pay some graduate students/postdocs to solve some math problems. The publicity of solving such benchmark would definitely raise the valuation so it would 100% be worth it for them...

Simple: I highly doubt they're willing to risk a scandal that would further tarnish their brand. It's still reeling from last year's drama, in addition to a spate of high-profile departures this year. Not to mention a few articles with insider sources that aren't exactly flattering.

You’re not thinking about the other side of the equation. If they win (becoming the first to excel at the benchmark), they potentially make billions. If they lose, they’ll be relegated to the dustbin of LLM history. Since there is an existential threat to the brand, there is almost nothing that isn’t worth risking to win. Risking a scandal to avoid irrelevance is an easy asymmetrical bet. Of course they would take the risk.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#88

Earlier quoted context omitted.

You could but most LLMs can't solve sudoku puzzles even though the training corpus already contains books on logic, constraint propagation, and state space exploration with backtracking.

the LLM just doesn't have enough compute I think probably step by step it could do it. anyways most leading LLMs can write a backtracking search program to do it and I think +tool use should be counted

[dead]

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#89
post #18

Earlier quoted context omitted.

Most humans can't solve these problems, so it's certainly possible to imagine a legitimate AGI that can't either.

AGI should be able to do anything the best humans can do. ASI is when it does everything better than the best humans.

Those thresholds look the same to me, personally.

An AI that can be onboarded to a random white collar job, and be interchangeably integrated into organisations, surely is AGI for all practical purposes, without eliminating the value of 100% of human experts.

Re: FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI

#90

Earlier quoted context omitted.

An AGI is often claimed to be a general purpose problem solver and these are exactly the types of problems that a general purpose problem solver would be able to solve if given access to a mathematical library. All existing LLMs have been trained on abstract mathematics and logic but it is obvious that they are incapable of abstract logical reasoning, e.g. solving sudoku puzzles.

99.9% of the population would not be able to solve these problems given a year and access to every piece of mathematical literature ever written (except the solutions to these problems of course). Saying that you need to solve these to be considered AGI is ridiculously strict.

[dead]
Post reply on HN