Live data from Hacker News

Ten advances in mathematics and theoretical computer science

openai.com

681–690 of 1000 posts

Re: Ten advances in mathematics and theoretical computer science

#681
post #636

Looking at this thread, I can see that a lot of technical people have ambivalent to negative feelings towards AI, but with each new generation, I become more and more convinced that they're missing out on something interesting. It is indeed true that all models are, at their core, predictors of what occurs next in a sequence. But I think it's worth exploring the implication of what that means. Because when fed tiny p…

> I can see that a lot of technical people have ambivalent to negative feelings towards AI My negative feelings towards AI are about energy use and inequalities, that kind of stuff. It undeniably works well, but whether or not it is better for society or the planet is a lot less clear.

Same. Even the most basic versions of LLM are magical to me. To encode meaning from language like that, and to then form relationships based on it, and use it to solve problems. It's an amazing technical achievement.

But I want to be reading about that from the comfort of a home, with a full belly.

Re: Ten advances in mathematics and theoretical computer science

#682

Can’t wait for this stuff to have quality of life increases for the average person. So far all I see is that AI has made owning a computer more expensive, made some jobs redundant, increased spam and distrust with questionable authenticity of content and of course made some Americans very rich.

It has had significant quality of life increases for me. I use LLMs for everything from:

- travel and restaurant recommendations. my last few outings have been entirely LLM-advised and they turned out excellent. LLMs seem to have ingested every single Google review, photo, and menu of every business on Earth and can answer very nuanced questions like "is the garlic chicken at garnished with coriander?"

- fitness, nutrition, accounting, therapy, medical, legal, immigration advice (sure it's not a real professional but you know what, it's pretty fucking close, and any capability gap is made up by having perfect two-way communication which you don't get when talking with a human)

- coding (work, side projects, personal tools, documentation & pricing questions, "review this code", etc).

- I start reading most articles with the prompt "Summarize this article: ". I just started a non-fiction book by pasting into Claude: "There are 12 chapters in the book . Can you give me a 2 sentence synopsis of each chapter?". It reduces the "activation energy" hump and screens if it's worth reading at all.

- I use the LLM in my Tesla for on-the-fly advice for parking and other things. You can simply ask "what's the best Boba place around here?" and it will give you a decent recommendation. You can also follow up with "does this place have ample parking?".

- I use the LLM in YouTube to summarize videos and ask specific questions and/or get timestamps to the parts I care about.

If your critical thinking skills are strong then LLM is a literal superpower.

Re: Ten advances in mathematics and theoretical computer science

#683
post #662

Earlier quoted context omitted.

this gets more nuanced because "the sigmoids won't save you": https://www.astralcodexten.com/p/the-sigmoids-wont-save-you

If the sigmoid is incorrect it's certainly more correct than the exponential. > https://www.astralcodexten.com/p/the-sigmoids-wont-save-you The conclusion of this article seems to be "you should give ai the benefit of the doubt against all reason". Barf

Isn’t the point more “it’s easy to fall into the trap to believe that predicting when the sigmoid is going to bend is possible and the right heuristic is to instead extrapolate locally”?

That aside, I’d question whether applying the Lindy effect in particular to something that’s not really a life expectancy but more a growth rate is credible… or perhaps a bit circular since it “assumes away” the ceiling.

Re: Ten advances in mathematics and theoretical computer science

#684
post #8

I'm enjoying learning about these hard problems, but this line about credit made me chuckle: > We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?

I'm not an expert at it myself, but my understanding is there are numerous ways to "cheat" in a Lean proof (via `sorry` and similar). They're taking responsibility for fully verifying that none of these cheats were used (and that the theorem statements themselves were all correctly formalized.)

Even beyond cheating with sorries or kernel bugs, the lean encoded theorems (or specifications) must be checked by humans to see if they truly mirror the real theorem authentically.

Re: Ten advances in mathematics and theoretical computer science

#686

Earlier quoted context omitted.

Proof? In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.

For customer support I don't think models have gotten better since gpt-4.1. The class of small models, with limited to no reasoning, that need to handle a complex issue with a touch of empathy, has not improved much. I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following max…

I understand the point (I don't agree with it; tool calling has gotten much better/reliable and that is very important for customer support) but consider: If you can get same for a lot less, that's an improvement. If we found a way to supply fresh water and electricity for -90% cost after 2 years, that would be fantastic.

You can do many more things, when stuff is cheaper, even if the stuff were otherwise unchanged.

Re: Ten advances in mathematics and theoretical computer science

#687
post #676

Earlier quoted context omitted.

So your answer is: ignore the progress, it’s not really happening, actually it’s getting worse. That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.

It sounds like they are making a clear argument: models are getting worse for certain domains even while they are getting better at others. I don't know if I agree with that but it doesn't seem like an irrational claim and does seem credible to me.

I just don’t think it’s true, or is significant enough to matter to the direction of travel of AI even if there were something to it. It’s another cope post being lobbed at the idea of AI going somewhere and I’m sick of them.

Re: Ten advances in mathematics and theoretical computer science

#688

People argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps delivering more and more impressive results. The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn, but I’ve noticed Fable to be quite a big step up there. How about politics? Will we develop new ways to let…

Another interesting question is why the frontier labs are piling on pure maths, which has little direct economic value compared to something like law or improving the efficiency of their own models? How much OpenAI and Anthropic are paying to serve these models for ordinary users is the elephant in the room. A cynical take is that the frontier labs are trying their best to pump up their pre-IPO valuation through flashy headlines.

Re: Ten advances in mathematics and theoretical computer science

#689

Earlier quoted context omitted.

> but I’ve noticed Fable to be quite a big step up there what did you notice ?

Fable is much better at handling nuance. Opus/GPT 5.6 Sol are much more likely to miss the point you are trying to make, emphasise the wrong thing, exaggerate the importance of unimportant details, or introduce contradictions. That said, Fable is still not a great writer, largely driven by it not knowing what it should exclude, and it still having the usual LLM-isms. But it’s better.

That's really what got everyone hooked in the first place.

5.6 Sol is great but there's a depth to the understanding that Fable exhibits that's unique to it currently.

Can I truly quantify this? I don't think so. Just that I spend a ton of time with various models and a certain point it's just a personal impression or a gut feeling.

In the days after Fable first came out I increased the amount of parallel planning of tasks that I was doing by 2-3x because it felt like I didn't need to be paranoid due to that handling of nuance.

Re: Ten advances in mathematics and theoretical computer science

#690
post #672

Earlier quoted context omitted.

It’s a marketing. They are a sham company. If this article was by Scientific American or something it would be worth a lot more. They are literally trying to keep the hype train on track. Also on HN front page today: AI's debt binge can't last, hidden borrowing reaches $1.65T (fortune.com) https://news.ycombinator.com/item?id=49160699

It is marketing, they are a shady company, and yet, if someone had access to these results before today's modern AI tools, they could get tenure at any university in the world.

As others have said, it's hard to know how significant these results are without more transparency around the methods used to obtain them.
Post reply on HN