Live data from Hacker News

GPT-5.6 used a prompt to close a 30-year gap in convex optimization

old.reddit.com

371–380 of 414 posts

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#371

Earlier quoted context omitted.

Surgeons too now? Another avenue for human skill destroyed. I actually want to exit. I want to live in a society where humans flourish not AI. Actually one just needs to walk the streets of Japan and compare that to US. Tokyo has hundreds of small shops with humans doing specific niche stuff, perfecting their arts. That’s all so beautiful. America has massive warehouses and supermarkets, with completely uninterested…

What you write is what's profoundly depressing. The only way you see people flourishing is through coercion, by society or by the world. You'd explicitly rather prolong this codependency for your aesthetic preferences even, as per your own words. We're already just circus monkeys in your world. You cannot fathom people pursuing self-improvement because e.g. developing or being capable is inherently enjoyable. You onl…

Yes, not everyone will agree with me. Maybe some people just want to consume AI abundance forever and live on UBI, like Altman’s wet dream.

But there will be some group of people, the people who wanted to be skilled tailors, skilled wood workers and apply them towards others before industrialization ruined them and of course the best mathematicians, the best programmers, before AI revolution will ruin them. We want no part of this technology, and maybe in some distant future we can organize and live separately from the consumerist dystopian hellscape that the world around us turns into.

I’m not charting a course for the world to follow, but I suspect every OpenAI employee or Sam Altman have no desire to live their entire life on UBI doing nothing of any importance. This for them is their pursuit of greatness, only in the result they’re destroying it for large swaths of humanity that live around them. I don’t want any part of it, I refuse to be a WallE human managed by AI Sam’s company designs to be put in a state of perpetual limbo somewhere.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#372

Earlier quoted context omitted.

What was the first DeepSeek moment? (genuine question, I'm out of the loop on what you mean)

I guess when cheap Chinese model was very close to SoTA models

... and now another one when a (not so cheap tbh) Chinese model matched OpenAI GPT Sol and Anthropic Fable/GPT 4.8 (some scores below, some scores above) ...

Oh and exceeded Google's best internal model capability enough to have them cancel the next Gemini release. Probably. But I'm 95% certain that's exactly what happened. This is the first time I really fear for the future of Google, because OpenAI matched or exceeded the experience of Google search ... and now a Chinese company matched OpenAI ... followed by Google failing to catch up. That really made me take pause.

It indicates the the "secret sauce" of OpenAI and Anthropic is ... nothing, really. Or none of it matters, except access to the hardware needed for 1.5 trillion parameter models, which Chinese firms now have as well. It means there is no "AI takeoff" where other firms can't match OpenAI/Anthropic models. And Google, inventors of transformers just missed the takeoff. There's more secret sauce, but clearly a great engineering firm can find it in less than 3 months.

The fact that this Chinese model has similar cost to OpenAI and ballpark cost compared to Anthropic while we can be quite sure they're minimizing the cost (that's what China is doing everywhere else) also raises questions about the true cost of serving comparable models and with that about the future profitability of OpenAI and Anthropic.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#373

Earlier quoted context omitted.

Sometimes I read a comment on HN that is so advanced that it's just as readable to me as Greek. Love reading it just to see someone work though!

> so advanced that it's just as readable to me as Greek I used to feel this way about statistics. The language and terms are hard to understand and many of the formulas are taught as "just memorize this" instead of building up from first principles. But then I started using statistics to analyze something I cared a lot about (paintball) and I quickly realized it's like learning anything new: - there is jargon - and c…

This is my exact experience right now trying to wade through the research on psychometrics and skill/knowledge assessment design. It’s mostly just applied statistics but like all such fields, over decades of specialization it acquired its own jargon for abstractions that are quickly recognizable to anyone with a sufficiently developed nose for modern mathematics. But you still have to wade through all the definitions to make those connections before you actually understand everything.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#374

Earlier quoted context omitted.

Sometimes I read a comment on HN that is so advanced that it's just as readable to me as Greek. Love reading it just to see someone work though!

Not to diminish the comment, but most things are not as complex as they sound when phrased in everyday language or sound much more complex than they are when phrased in technical language. Technical language is a tool that allows insiders to say less and refer to more, and to be specific, but it's just a tool. Most things can be described in accessible ways. I think you'd be surprised at what you could understand and…

Yeah, heck, whenever an LLM puts my thoughts and intuition into words, it sounds really complex as well.

(FWIW I have an issue with producing words, rather than recognition. I do have the intuition I just lack the labels for it.)

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#375
post #249

Earlier quoted context omitted.

I'm from Denmark and I've been an external examiner for various CS educations for the previous 13 years now. Some of them teach you a lot about how the hardware works, others mainly teach you design patterns. Five years ago the latter was in high demand, because a lot of software development frankly doesn't need computer science (until it does). Now there is almost no demand for them.

Honestly, i've received a formal MSc education in the hardware aspects, including for designing embedded electronics products. Spent the most part ofy career in the software industry designing enterprise software and feel like i never needed to use them, except maybe early in my career when i was reviewing tech stacks and determined that .NET would be among the winning horses, precisely because it'd take care of that…

> However, having a high affinity with hardware is not a driver / computer science of hiring decisions from what i can see in the enterprise software world

I think the way I worded it was maybe a little too close to just being about hardware, because performance do matter a lot in the energy industry. I do think it applies to SWE in general. You mention .NET and I've met C# developers with years of experience who couldn't tell you the difference between IEnumerable and IQueryable. I've met even more experienced Python developers who don't know what a generator is. Stuff like that, not having knowledge of the tools they use. I guess you could argue that those are bad developers, but I don't personally think that has been the case for most of them. Still, you'd rather have someone who thinks about these things rather than eventually using batches once they run into memory issues.

I also think these changes are appearing faster in non SWE enterprise. As you said, product owners who are AI explorative (for the lack of a better word) are rock stars. We see this a lot in our finance and risk departments, where domain experts now write fairly decent software with AI. My team has build them tools so that they build things the same way, use the same developer setups and pull the pre-approved external packages or are offered alternative ways of doing things. A few years ago this would've been done by these domain experts "ordering" the software they needed from our SWE team, and if they hadn't already been mostly laid off due to Putin's invasion of Ukraine changing the markets, I belive they would've been now because of AI.

Because frankly, a lot of the software that gets produced in these areas, don't need computer science, until it does, and the domain experts can make the software they need so much faster than before by vibe coding it. From my perspective it's not that much of a difference in the quality of the code that gets produced. I also had to help with performance and security when we had more software engineers on staff. Though now I do it more through writing and distributing AI agent applications rather than writing a C binary or optimising the code directly.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#376

Earlier quoted context omitted.

Mathematics is a human-designed game that involves rearranging symbols.

That view is incredibly reductionist. It really is an efficient encoding of how nature behaves. It might be a human construct, but given how best it allows to understand nature (through principles of physics), it is uncanny to be any different from the language of nature. Reminds me of Wigner's Unreasonable effectiveness of mathematics in natural sciences [0]. [0]: https://en.wikipedia.org/wiki/The_Unreasonable_Effec…

Nature has no language. Language is a human tool to imperfectly describe nature. Putting language before nature is magical thinking (also known as Platonism).

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#377

Earlier quoted context omitted.

I'd push back on this. Most of the core optimization techniques (eg, ADAM, stochastic gradient descent) are straight out of the convex optimization literature. Generally you need to use optimizers that work well on convex objectives because near minimizers, functions tend to be convex. (Proof by contradiction: a non-convex point has a strict descent direction.) The fact that neural networks are highly nonconvex has e…

I'd say it's going to be very hard to come up with a method that works on general nonconvex functions while not working on convex functions

It's not a matter of whether the theory "works"; it's a matter of whether one is asking the right questions. Convex optimization studies how quickly an optimizer can reach the optimum. In the non-convex case, there are many basins containing their own local minima. The more sensible questions there are "which basin is it likely to go into?" and "how do I steer it to go where I want?". Global convergence rates are largely irrelevant by comparison.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#378
post #267

Earlier quoted context omitted.

Yours' "not as efficient" in [2] means that, sometimes, ADAM "does not work." Look at figure 2, ADAM literally does not work in the case of "true model."

Yes, apologies, I didn't read the articles you linked before posting this. I did update the comment. I don't think this changes the point, which is that most optimization methods used in AI owe a substantial intellectual debt to convex optimization theory.

I love convex optimization and there are a few SciML projects I am on where I really need results from there. But in AI research with deep neural networks, it's become a liability, because people will just not let go. I'm getting tired of reviewing convex optimization theory papers in ML conferences that are still trying to wave away the obvious issues with their application to deep learning. It's harsh, but I do feel we can only start talking about an intellectual debt once that stops being the case.

Re: GPT-5.6 used a prompt to close a 30-year gap in convex optimization

#379

Earlier quoted context omitted.

Well that’s a problem of incentives. Why would a manager outsource their own job to an AI?

It's not a problem of incentives. Every executive wants to inject LLMs everywhere these days. If they haven't somewhere it means that it does not work.

Every executive wants to inject LLMs downstream from their own job. As you get away from execution and toward management the incentive to innovate gets less and less as management at the same level.
Post reply on HN