Earlier quoted context omitted.
It's not at all a joke ... that's a severe misunderstanding of the context.
There is no evidence that I can find for the claim "a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle." As I currently understand it, all we know is that: - a mathematician produced a Lean-verified counterexample to the Collatz conjecture, demonstrating a bug in the kern…
Ten advances in mathematics and theoretical computer science
611–620 of 1001 posts
Re: Ten advances in mathematics and theoretical computer science
#612Earlier quoted context omitted.
Well I don't typically side with GM, but playing devil's advocate: 1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet? 2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer? 3. Not wrong. 4. I think he'd probably pull you up on 'bug free' - I don't think that frontier model…
1 is wrong. If I tell Codex + GPT-5.6 to do it now, it will figure out how to do it. If it would need to extract audio and run a speech model on it, it will find one, set it up, and run without my help.
An LLM could theoretically try to earn some money and pay a human to do most of these tasks but it's not the point of the exercise.
Re: Ten advances in mathematics and theoretical computer science
#613Earlier quoted context omitted.
you are working on coding. they are working on things like "creative writing" remember that gpt 4o was popular among those who had ai as a romantic partnet?
> remember that gpt 4o was popular among those who had ai as a romantic partner I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.
Re: Ten advances in mathematics and theoretical computer science
#614People argue whether we are at y-5, y, or y+5, meanwhile we seem to be on a y=2^x exponential that keeps delivering more and more impressive results. The most interesting question to me is what will be consumed by the exponential like math seems to be undergoing, and what won’t. Writing has been quite stubborn, but I’ve noticed Fable to be quite a big step up there. How about politics? Will we develop new ways to let…
> but I’ve noticed Fable to be quite a big step up there what did you notice ?
That said, Fable is still not a great writer, largely driven by it not knowing what it should exclude, and it still having the usual LLM-isms. But it’s better.
Re: Ten advances in mathematics and theoretical computer science
#615Earlier quoted context omitted.
Well I don't typically side with GM, but playing devil's advocate: 1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet? 2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer? 3. Not wrong. 4. I think he'd probably pull you up on 'bug free' - I don't think that frontier model…
4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P
Re: Ten advances in mathematics and theoretical computer science
#616Not being an expert in any of the fields OpenAI has "advanced" I don't want to prematurely downplay the significance of this contribution. However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing. It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus,…
You’ve received the expert answer several times. You just don’t seem to like the answer.
Re: Ten advances in mathematics and theoretical computer science
#617Earlier quoted context omitted.
Ask DraftKings?
You'd need this argument to be a lot more concrete as to why AI is like gambling.
Perhaps you can ask Claude to explain it to you.
Re: Ten advances in mathematics and theoretical computer science
#618Earlier quoted context omitted.
I can deal with apathy, that’s the norm. What bothers me are all the people who think they can suppress AI by talking it down. That’s what’s counterproductive, just pretend the problem doesn’t exist. Tell other people it doesn’t exist either. I get it, it’s threatening socially, economically, maybe existentially. It’s also not going away.
So you think it’s a potentially existential threat but are bothered by people who maybe want to suppress it… Hopefully you acknowledge there is a bit of lack of self awareness here eh?
Re: Ten advances in mathematics and theoretical computer science
#619Not being an expert in any of the fields OpenAI has "advanced" I don't want to prematurely downplay the significance of this contribution. However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing. It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus,…
Also on HN front page today: AI's debt binge can't last, hidden borrowing reaches $1.65T (fortune.com)
Re: Ten advances in mathematics and theoretical computer science
#620Earlier quoted context omitted.
> Whilst current models can't 'intuit' and come up with conjectures People keep saying this. Why? Surely the AI can complete the prompt “Generate new research questions based on these observations”? When I read the reasoning traces of coding models they are constantly asking themselves questions and attempting to answer them.
I like the illustration that the models are working on a convex hull of known information. Filling gaps with linear combinations of known facts and results. They can't exit the hull until the "intuition" starts spawning points outside the convex hull.
The extent to which they are able to do this is the more interesting question!