Live data from Hacker News

Ten advances in mathematics and theoretical computer science

openai.com

621–630 of 1001 posts

Re: Ten advances in mathematics and theoretical computer science

#621
post #96

Replace philosophers for mathematicians and Douglas Adams was spot on again. Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this. -- "Yes we are," insisted Majikthise. "We are…

Ahh, but you missed the continuation, where they get to the heart of the matter: money.

"Excuse me, We demand rigidly defined areas of doubt and uncertainty!"

DT: Might I make an observation at this point?

MT: You keep out of this metal nose.

VF: We demand that that machine not be allowed to think about this problem!

DT: If I might make an observation…

MT: We’ll go on strike!

VF: That’s right. You’ll have a national philosopher’s strike on your hands.

DT: Who will that inconvenience?

MT: Never you mind who it’ll inconvenience you box of black legging binary bits! It’ll hurt, buster! It’ll hurt!

DT: [Booming] If I might make an observation …

“All I wanted to say,” bellowed the computer, “is that my circuits are now irrevocably committed to calculating the answer to the Ultimate Question of Life, the Universe, and Everything.” He paused and satisfied himself that he now had everyone’s attention, before continuing more quietly. “But the program will take me a little while to run.”

Fook glanced impatiently at his watch.

“How long?” he said.

“Seven and a half million years,” said Deep Thought.

Lunkwill and Fook blinked at each other.

“Seven and a half million years!” they cried in chorus.

“Yes,” declaimed Deep Thought, “I said I’d have to think about it, didn’t I? And it occurs to me that running a program like this is bound to create an enormous amount of popular publicity for the whole are of philosophy in general. Everyone’s going to have their own theories about what answer I’m eventually going to come up with, and who better, to capitalize on that media market than you yourselves? So long as you can keep disagreeing with each other violently enough and maligning each other in the popular press, and so long as you have clever agents, you can keep yourselves on the gravy train for life. How does that sound?”

The two philosophers gaped at him.

“Bloody hell,” said Majikthise, “now that is what I call thinking. Here, Vroomfondel, why do we never think of things like that?”

“Dunno,” said Vroomfondel in an awed whisper; “think our brains must be too highly trained, Majikthise.”

So saying, they turned on their heels and walked out of the door and into a life-style beyond their wildest dreams.”

Re: Ten advances in mathematics and theoretical computer science

#622
post #342

Earlier quoted context omitted.

The Wozniak test has it with a robot body going into a house, finding a coffee maker and making a cup. I guess you can vary the rules as you like.

That's a good test for a robot, not for AGI. AGI should test only intelligence and should not require limbs.

Okay we can make a completely digital environment for it to make coffee in then. It would still fail unless you let it randomly try every combination potentially thousands of times until it stumbles upon the right path. It doesn't take intelligence to read off a recipe, it does take intelligence to read a recipe, understand it, adapt it to your specific tools and materials on hand which may differ from the recipe, and then actually still accomplish it in the first or maybe second try.

Re: Ten advances in mathematics and theoretical computer science

#623

People argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps delivering more and more impressive results. The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn, but I’ve noticed Fable to be quite a big step up there. How about politics? Will we develop new ways to let…

There are many math problems that are simply puzzles: intellectually interesting but nothing worth of value depends on it. To me it would be more impressive if we could define hard problems that need to be solved up front and see how the models deal with that.

The results OpenAI demonstrated are impressive, but it also looks like they threw a lot of compute at it just to get results. How many tokens did they waste on problems they couldn't solve? Applying inference infrastructure on a large number of math problems at scale we haven't seen before to me doesn't demonstrate an exponential curve in model abilities.

Re: Ten advances in mathematics and theoretical computer science

#625

Earlier quoted context omitted.

UBI probably is the positive outcome, though it may not seem like it to begin with. Initially it will likely be stigmatised and under-resourced, but as a larger proportion of people move out of work and onto UBI that stigma will drop and the resources should grow. Eventually UBI will be the norm, and if the living standards of a person on UBI is as good as yours or mine today, that will be an enormous win for everyon…

> Eventually UBI will be the norm I see comments like this tossed around a lot, but what makes you say this? Don't you think its more likely that most people end up in poverty?

No, I think people will demand to not live in poverty. Capitalism allows people to escape poverty by personal effort. An AI future won’t allow it by any means other than collective effort, so that’s what we’ll do.

Re: Ten advances in mathematics and theoretical computer science

#626

Earlier quoted context omitted.

Not everyone works for Evil Corp. I work in the public sector and my work supports public health and safety initiatives. AI has allowed my team to get much more done than we would have otherwise which improves the quality of life of the people in my community. So I would like to counter your cynicism with a “YMMV” depending on who you work for.

Are you just automating lab reporting and results faster? That doesn’t exactly improve the health of others, a faster lab does not cure an ailment or provide a better cure. Like cool my lung xray only took minutes to determine if I have a lesion instead of a week or a few days, but I still have cancer.

For some people, getting lung cancer reported a week earlier will save their lives.

More importantly, if you can screen for cancer in a way that takes minutes instead of a week, imagine how accessible this technology will become.

Re: Ten advances in mathematics and theoretical computer science

#627
post #381

Earlier quoted context omitted.

[flagged]

What makes you think separate nations would effectively cooperate in this way?

Might as well try, no? Or do you think it's better we just risk it? I mean it's not like the West has any hard or soft power to encourage cooperation.

Re: Ten advances in mathematics and theoretical computer science

#628
post #306
post #186

Earlier quoted context omitted.

"Breakthrough research" can be defined (in the citation record) as research that both (1) becomes highly cited, and (2) brings together citation chains that were previously not showing up together. Mundane incremental research is cobbled from existing citations that already appear nearby in the record. Basically, innovative research is a measure of bridging thought and domains that were previously not bridged. It's q…

I'm not arguing that this isn't innovative or worthy of publication. Basically any result that moves the needle meets those criteria. I'm interested in how the results that OpenAI has published here differs from finding optimality solutions for incredibly niche optimization problems by throwing the problem in an enormous solver.

They differ in that there hasn't been a solver you could have thrown them at. I guess you could argue their harness + LLM setup is a "solver", but the approach is so different from what we used that word for in the past that I don't think it would be appropriate.

Re: Ten advances in mathematics and theoretical computer science

#629

Gary Marcus' has a good take on this: https://garymarcus.substack.com/p/openais-amazing-but-vastly... https://garymarcus.substack.com/p/two-critical-updates-re-as... Not that there isn't something interesting in here, but lets be clear that we don't have enough information to evaluate this properly. And as always with these labs, BS takes a lot more energy to refute than it does to spread.

I find Marcus on this, something approaching sophistry and rhetorical showmanship in service of maintaining an ideological position, for reasons unrelated to the nominal intellectual clarity. To sharpen that, I think he's (obviously) interested in maintaining his own brand as "thought leader" and this necessitates de rigeur defense of particular postures. Sometimes this is easy because the facts warrant it; other tim…

[deleted]

Re: Ten advances in mathematics and theoretical computer science

#630

Earlier quoted context omitted.

The models are frequently getting worse at items that they aren’t being benchmarked for — and that’s happening more and more over time! Other people in other fields aren’t idiots, they are accurately perceiving the fact that these models are being hyper optimized for our industry, and are becoming less capable in other domains over time. Models of the same scale are massively worse at writing a broad variety of style…

Proof? In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.

For customer support I don't think models have gotten better since gpt-4.1. The class of small models, with limited to no reasoning, that need to handle a complex issue with a touch of empathy, has not improved much.

I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following maximizing models seem to make worse free-form agents, but they're really all that some domains need.

Post reply on HN