Live data from Hacker News

Ten advances in mathematics and theoretical computer science

openai.com

681–690 of 1001 posts

Re: Ten advances in mathematics and theoretical computer science

#681
post #660

Earlier quoted context omitted.

this gets more nuanced because "the sigmoids won't save you": https://www.astralcodexten.com/p/the-sigmoids-wont-save-you

If the sigmoid is incorrect it's certainly more correct than the exponential. > https://www.astralcodexten.com/p/the-sigmoids-wont-save-you The conclusion of this article seems to be "you should give ai the benefit of the doubt against all reason". Barf

Isn’t the point more “it’s easy to fall into the trap to believe that predicting when the sigmoid is going to bend is possible and the right heuristic is to instead extrapolate locally”?

That aside, I’d question whether applying the Lindy effect in particular to something that’s not really a life expectancy but more a growth rate is credible… or perhaps a bit circular since it “assumes away” the ceiling.

Re: Ten advances in mathematics and theoretical computer science

#682
post #8

I'm enjoying learning about these hard problems, but this line about credit made me chuckle: > We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?

I'm not an expert at it myself, but my understanding is there are numerous ways to "cheat" in a Lean proof (via `sorry` and similar). They're taking responsibility for fully verifying that none of these cheats were used (and that the theorem statements themselves were all correctly formalized.)

Even beyond cheating with sorries or kernel bugs, the lean encoded theorems (or specifications) must be checked by humans to see if they truly mirror the real theorem authentically.

Re: Ten advances in mathematics and theoretical computer science

#684

Earlier quoted context omitted.

Proof? In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.

For customer support I don't think models have gotten better since gpt-4.1. The class of small models, with limited to no reasoning, that need to handle a complex issue with a touch of empathy, has not improved much. I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following max…

I understand the point (I don't agree with it; tool calling has gotten much better/reliable and that is very important for customer support) but consider: If you can get same for a lot less, that's an improvement. If we found a way to supply fresh water and electricity for -90% cost after 2 years, that would be fantastic.

You can do many more things, when stuff is cheaper, even if the stuff were otherwise unchanged.

Re: Ten advances in mathematics and theoretical computer science

#685
post #674

Earlier quoted context omitted.

So your answer is: ignore the progress, it’s not really happening, actually it’s getting worse. That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.

It sounds like they are making a clear argument: models are getting worse for certain domains even while they are getting better at others. I don't know if I agree with that but it doesn't seem like an irrational claim and does seem credible to me.

I just don’t think it’s true, or is significant enough to matter to the direction of travel of AI even if there were something to it. It’s another cope post being lobbed at the idea of AI going somewhere and I’m sick of them.

Re: Ten advances in mathematics and theoretical computer science

#686

People argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps delivering more and more impressive results. The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn, but I’ve noticed Fable to be quite a big step up there. How about politics? Will we develop new ways to let…

Another interesting question is why the frontier labs are piling on pure maths, which has little direct economic value compared to something like law or improving the efficiency of their own models? How much OpenAI and Anthropic are paying to serve these models for ordinary users is the elephant in the room. A cynical take is that the frontier labs are trying their best to pump up their pre-IPO valuation through flashy headlines.

Re: Ten advances in mathematics and theoretical computer science

#687

Earlier quoted context omitted.

> but I’ve noticed Fable to be quite a big step up there what did you notice ?

Fable is much better at handling nuance. Opus/GPT 5.6 Sol are much more likely to miss the point you are trying to make, emphasise the wrong thing, exaggerate the importance of unimportant details, or introduce contradictions. That said, Fable is still not a great writer, largely driven by it not knowing what it should exclude, and it still having the usual LLM-isms. But it’s better.

That's really what got everyone hooked in the first place.

5.6 Sol is great but there's a depth to the understanding that Fable exhibits that's unique to it currently.

Can I truly quantify this? I don't think so. Just that I spend a ton of time with various models and a certain point it's just a personal impression or a gut feeling.

In the days after Fable first came out I increased the amount of parallel planning of tasks that I was doing by 2-3x because it felt like I didn't need to be paranoid due to that handling of nuance.

Re: Ten advances in mathematics and theoretical computer science

#688
post #670

Earlier quoted context omitted.

It’s a marketing. They are a sham company. If this article was by Scientific American or something it would be worth a lot more. They are literally trying to keep the hype train on track. Also on HN front page today: AI's debt binge can't last, hidden borrowing reaches $1.65T (fortune.com) https://news.ycombinator.com/item?id=49160699

It is marketing, they are a shady company, and yet, if someone had access to these results before today's modern AI tools, they could get tenure at any university in the world.

As others have said, it's hard to know how significant these results are without more transparency around the methods used to obtain them.

Re: Ten advances in mathematics and theoretical computer science

#689

Earlier quoted context omitted.

> but I’ve noticed Fable to be quite a big step up there what did you notice ?

I've noticed that out of all LLMs I've ever used that Fable is the MOST LLM; the text it produces is abomination. It's impressive how much I hate it. It is such an awful writer - it assumes the reader has zero context and therefore gives every single bit of context and detail - which is nice if you're writing a legal document I suppose. But it uses, niche, $10 words to describe every facet of everything it's discussi…

Just praised Fable in another comment but what you're saying is also insanely true.

I literally roll my eyes and cringe quite often at its output pretty much daily.

I don't like to overload my sessions with skills but I've been using a "write-normal" skill I made just to have it rewrite outputs that particularly piss me off.

https://gist.github.com/alasano/1c734fa055231a5defcfd213217e...

I'm sure there's a million of these skills out there, but this one is tailored to the stuff that makes me mad in particular.

Re: Ten advances in mathematics and theoretical computer science

#690
post #660

Earlier quoted context omitted.

this gets more nuanced because "the sigmoids won't save you": https://www.astralcodexten.com/p/the-sigmoids-wont-save-you

If the sigmoid is incorrect it's certainly more correct than the exponential. > https://www.astralcodexten.com/p/the-sigmoids-wont-save-you The conclusion of this article seems to be "you should give ai the benefit of the doubt against all reason". Barf

The author of that post is a prominent Bay Area "rationalist," who have had a quasi-theistic relationship with the concept of all-powerful AIs for a couple decades now.
Post reply on HN