Live data from Hacker News

GPT-5.6 used a prompt to close a 30-year gap in convex optimization

old.reddit.com

131–140 of 414 posts

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#131

Earlier quoted context omitted.

It's still clear that LLMs lack spatial reasoning, either in the concrete or abstract, and while that sort of reasoning has been downplayed by academia for at least a century it is fundamental to technology and industry. (And many would say for science and mathematics too). They will, however, get there as well either directly or as interfaces to models that do, and your core point stands.

Is there any proof that they are not good at special reasoning? Arc agi 1 and 2 are saturated.

ARC AGI 3 is much better designed and harder, perfectly completable by a human in a couple minutes.

Only a fraction of the games can be solved by Sol, generally at sub-human efficiency in terms of turns, AND at a cost of >$10,000 per game.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#132

Genuine question: If you still or did think LLMs are just stochastic parrots that just summarize everything and have no form of creativity, what do you think after seeing results like this? I'm very curious how people reconcile their fear/hatred of AI with actual objective reality. This is actually what interests me most about the whole AI thing. How we tell ourselves what we tell ourselves.

I had to create an account to respond to this because I am quite convinced these math problems they are "solving" are pure marketing. Why is it only GPT doing this, why not Claude? Why does Terrance Tao do marketing for OpenAI? I suspect OpenAI has hired math researchers to solve obscure problems and put them in their training set, purely for marketing reasons.

There was a good comment on the Pelican bicycle svg yesterday about how these models aren't getting much better beyond what the companies focus training them on. I think that's what's happening in this case too, they probably put this in the training set.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#133
post #19

Earlier quoted context omitted.

That doesn't make any sense; you can't have one LLM to read your mind to prompt another LLM.

> you can't have one LLM to read your mind to prompt another LLM I’m excited to inform you that we as a species have developed a particularly useful facility known as Language which these LLM tools are evidently rather handy at wielding. This facility is particularly useful in this context when it takes the form of “dialog” or “questioning”, which can be used to propagate abstract ideas by means of mutually-feedback-…

so, promoting?

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#134
post #14

Crazy how intelligence is cheap, efficient and commonplace now. We humans better refocusing our energy on our core values/principles, given most of our skills are becoming irrelevant

Intelligence on its own is not very useful though. We put it on a pedestal because it creates huge potential when paired with other things, wisdom, discipline, empathy, but on its own?

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#135

Genuine question: If you still or did think LLMs are just stochastic parrots that just summarize everything and have no form of creativity, what do you think after seeing results like this? I'm very curious how people reconcile their fear/hatred of AI with actual objective reality. This is actually what interests me most about the whole AI thing. How we tell ourselves what we tell ourselves.

I hold my stance that LLMs are stochastic parrots. Making the parrots ever more complex and training on ever more data produced by intelligent, creative beings may make them more useful or convincing but does at no point give rise to intelligence or creativity.

With such high standards, most HN commenters also do not have intelligence nor creativity. I don’t think we can set the bar that high.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#136

Earlier quoted context omitted.

It's still clear that LLMs lack spatial reasoning, either in the concrete or abstract, and while that sort of reasoning has been downplayed by academia for at least a century it is fundamental to technology and industry. (And many would say for science and mathematics too). They will, however, get there as well either directly or as interfaces to models that do, and your core point stands.

"Lack" isn't the right word. "Lacking" is more like it. If there was a deep fundamental inability, we wouldn't see things like newer generations of LLMs consistently improving on ARC-AGI series (heavy spatial reasoning loading) and SimpleBench (a lot of commonsense + spatial reasoning components). In a way, it's a surprise that LLMs, notoriously lacking any sort of embodied experience, can even get this close to huma…

> "Lack" isn't the right word. "Lacking" is more like it.

Yeah, that's fair.

> My takeaway is that text is a far richer modality than anyone has expected - and that high end LLMs are often sharp and flexible enough to recognize their weak points and substitute their strengths. I.e. all the LLMs implementing A* to optimally solve pathfinding in ARC-AGI-3 tasks, often unprompted.

I agree and disagree with this. I think we've learned a lot of humans are more text based than we thought, but conversely I'm not persuaded what non-textual task reasoning LLMs are doing is necessarily text based, just that models have grown large enough for other reasoning modes to conceivably be hiding in the parameter space.

As I mentioned elsewhere, like many others I find LLMs work entirely by example, and reaching for A* when pathfinding is the single obvious thing to do. In cases where the magic key word is not mentioned and the problem cannot be identified as "pathfinding" (or some other trigger with a highly specific widely documented solution) they will struggle, yet the moment the trigger is hit they get there very fast. This is why prompting remains such an art form.

Fable is the first one I've encountered that is capable of serious open ended 3D programming in ways that suggest it has some grasp of the spatial aspects of the problem (not merely symbolic manipulation of the vectors etc.), but it still misses optimization opportunities a human will find glaringly obvious based on spatially predictable bounds etc.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#137
post #27

Earlier quoted context omitted.

Once we figure out the pesky problem of how we're going to pay for housing, food, and healthcare.

I think the big names behind the AI companies already have that problem solved. A lot of people probably won't like the solution very much though.

Yes, they have a final solution for all of us.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#138
post #34

Earlier quoted context omitted.

Oh brother AI hasn’t even taken the class of jobs associated with customer service lmao

Uh.... Have you ever called customer service lately?

Or indeed 20 years ago when "press 1 for foo, press 2 for bar" was already a thing.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#139

> I don't think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hanging (you know what I mean) fruit. We'll be needed for problems where actual novel approaches are needed. I wonder how this compares to what we see happening with "juniors" in software development? In math research, do you also get the training for the professio…

Around here AI isn't really more of a threat to juniors than it is to seniors. It's a threat to the people who have been taught "recipies" rather than applied computer science. You can have excellent seniors who can do TDD, DRY, SOLID and so on, who also happen to have no idea what a L1 cache miss is. The current AI models know all of those things, but they struggle applying them correctly without someone piloting th…

Unless you’re claiming that AIs will suddenly (and very soon) stop improving, they are obviously a threat to everyone’s job.

Calling notable conjectures that have been open for decades “low-hanging fruit” is an act of desperation. Most professional mathematicians couldn’t have proved those conjectures if their lives depended on it.

Post reply on HN