Live data from Hacker News

Ten advances in mathematics and theoretical computer science

openai.com

611–620 of 1001 posts

Re: Ten advances in mathematics and theoretical computer science

#611
post #83
post #78

Earlier quoted context omitted.

It's not at all a joke ... that's a severe misunderstanding of the context.

There is no evidence that I can find for the claim "a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle." As I currently understand it, all we know is that: - a mathematician produced a Lean-verified counterexample to the Collatz conjecture, demonstrating a bug in the kern…

Indeed. It seems to me much more likely that the AI was directed to look for bugs in Lean, found one, and then it was directed to write a proof specifically targeting the bug.

Re: Ten advances in mathematics and theoretical computer science

#612

Earlier quoted context omitted.

Well I don't typically side with GM, but playing devil's advocate: 1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet? 2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer? 3. Not wrong. 4. I think he'd probably pull you up on 'bug free' - I don't think that frontier model…

1 is wrong. If I tell Codex + GPT-5.6 to do it now, it will figure out how to do it. If it would need to extract audio and run a speech model on it, it will find one, set it up, and run without my help.

I'm not buying this. GM clearly was trying to set a benchmark for video comprehension, not tool usage. Video comprehension is required for many 'AGI tasks', especially robotics to work in real time.

An LLM could theoretically try to earn some money and pay a human to do most of these tasks but it's not the point of the exercise.

Re: Ten advances in mathematics and theoretical computer science

#613

Earlier quoted context omitted.

you are working on coding. they are working on things like "creative writing" remember that gpt 4o was popular among those who had ai as a romantic partnet?

> remember that gpt 4o was popular among those who had ai as a romantic partner I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.

Dude it's not a system prompt, it's the training.

Re: Ten advances in mathematics and theoretical computer science

#614

People argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps delivering more and more impressive results. The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn, but I’ve noticed Fable to be quite a big step up there. How about politics? Will we develop new ways to let…

> but I’ve noticed Fable to be quite a big step up there what did you notice ?

Fable is much better at handling nuance. Opus/GPT 5.6 Sol are much more likely to miss the point you are trying to make, emphasise the wrong thing, exaggerate the importance of unimportant details, or introduce contradictions.

That said, Fable is still not a great writer, largely driven by it not knowing what it should exclude, and it still having the usual LLM-isms. But it’s better.

Re: Ten advances in mathematics and theoretical computer science

#615

Earlier quoted context omitted.

Well I don't typically side with GM, but playing devil's advocate: 1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet? 2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer? 3. Not wrong. 4. I think he'd probably pull you up on 'bug free' - I don't think that frontier model…

4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P

The wording was 'reliably' though? I could just be splitting hairs on that one though to be honest.

Re: Ten advances in mathematics and theoretical computer science

#616
post #581
post #158

Not being an expert in any of the fields OpenAI has "advanced" I don't want to prematurely downplay the significance of this contribution. However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing. It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus,…

You’ve received the expert answer several times. You just don’t seem to like the answer.

Forgive me for taking everything salesmen say with a grain of salt.

Re: Ten advances in mathematics and theoretical computer science

#617
post #521

Earlier quoted context omitted.

Ask DraftKings?

You'd need this argument to be a lot more concrete as to why AI is like gambling.

This is not the argument. It's not a comparison to gambling but a comparison to something that does not materially improve a person's life. Economic expenditure does not equate to human benefit. This is the original argument, and the onus is on THAT person to explain why people spending for AI actually benefit, not the other way around.

Perhaps you can ask Claude to explain it to you.

Re: Ten advances in mathematics and theoretical computer science

#618
post #379

Earlier quoted context omitted.

I can deal with apathy, that’s the norm. What bothers me are all the people who think they can suppress AI by talking it down. That’s what’s counterproductive, just pretend the problem doesn’t exist. Tell other people it doesn’t exist either. I get it, it’s threatening socially, economically, maybe existentially. It’s also not going away.

So you think it’s a potentially existential threat but are bothered by people who maybe want to suppress it… Hopefully you acknowledge there is a bit of lack of self awareness here eh?

You seem to have misunderstood my point.

Re: Ten advances in mathematics and theoretical computer science

#619
post #158

Not being an expert in any of the fields OpenAI has "advanced" I don't want to prematurely downplay the significance of this contribution. However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing. It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus,…

It’s a marketing. They are a sham company. If this article was by Scientific American or something it would be worth a lot more. They are literally trying to keep the hype train on track.

Also on HN front page today: AI's debt binge can't last, hidden borrowing reaches $1.65T (fortune.com)

https://news.ycombinator.com/item?id=49160699

Re: Ten advances in mathematics and theoretical computer science

#620

Earlier quoted context omitted.

> Whilst current models can't 'intuit' and come up with conjectures People keep saying this. Why? Surely the AI can complete the prompt “Generate new research questions based on these observations”? When I read the reasoning traces of coding models they are constantly asking themselves questions and attempting to answer them.

I like the illustration that the models are working on a convex hull of known information. Filling gaps with linear combinations of known facts and results. They can't exit the hull until the "intuition" starts spawning points outside the convex hull.

Neural nets can extrapolate past their training data, and there is no reason to think LLMs don’t inherit this capability.

The extent to which they are able to do this is the more interesting question!

Post reply on HN