Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

401–410 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#401

I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am di…

I think that ultimately, the largest change brought on by LLMs will be due not to their intelligence, but to their tenacity.

If you had an infinite number of monkeys, each with a typewriter, one would eventually write Shakespeare. If you had an infinite number of college-educated interns, each with access to all the public records you can possibly get via FOIA, one would eventually get enough evidence to prove that a top politician is cheating on their partner, evidence which you could use to blackmail that politician.

You don't need that much intelligence to do that, you just need somebody who's willing to dedicate their life to knowing everything there is to know about that guy from Louisiana.

With humans, the amount of money you'd need to pay such a person just isn't worth the reward. With LLMs, it may very well be.

Re: A recent experience with ChatGPT 5.5 Pro

#402

Earlier quoted context omitted.

I think people's opinion of "marginal improvement" is based on their relative ability. A 2000 elo chess player is going to think the jump from 500 to 1000 is marginal. They're both floundering around not doing anything resembling common sense. A 1000 elo chess player is going to find the jump from 2000 to 2500 marginal. They're both playing far better moves for incomprehensible reasons, and the only reason you know t…

2024-2025 was filled with huge improvements. 2025-2026 has not been, outside of open source. The idea that we’re at the point where it’s superseded our ability to tell just makes no sense. I’ll be happy if we can get to a point where I don’t have to tell Claude not to tail every bash command or make a job that writes throughout instead of once at the end. I’ll be happy if “continue this interaction naturally, you are…

Claude in feb of 2025 was barely able to code. Sure, it could write you a nice function, it could even write you a complex 200-line algorithm, but give it a codebase, and it would quickly get overwhelmed.

Claude in feb of 2026? Still far from perfect, but there's definitely a huge improvement here.

Re: A recent experience with ChatGPT 5.5 Pro

#403

Earlier quoted context omitted.

We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro, not by familiar names that OpenAI is sneaking additional compute to behind the scenes (which is too conspiratorial of an explanation for my liking regardless).

> We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro Where's that? The stuff I've seen is from celebrities. Were those problems as hard as this one, or the ones that Tao posts about? Regardless.. what's the argument against more transparency here to just settle this kind of thing? > which is too conspiratorial of an explanation for my liking regardless OpenAI is…

> Where's that?

https://archive.ph/2w4fi

Re: A recent experience with ChatGPT 5.5 Pro

#404

This jives with what I've experienced in the brief time I had access to 5.5 Pro. It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. The downside (not noted in the…

> This jives with what I've experienced

Just as an fyi, the word you are looking for is jibes. Jive is something else entirely.

Re: A recent experience with ChatGPT 5.5 Pro

#405
post #334

Earlier quoted context omitted.

RL or no RL, AI cannot escape the distribution it's trained on. It's just that the labs will put so much into the distribution that we won't be able to tell the difference that easily, nor will it matter for most tasks. The reason AI does well on ARC-AGI-2 is because the labs created synthetic training data using similar puzzles.

Yes it can! That's the whole point of RL! it generates slightly out of distribution rollouts, and rewards good rollouts to change the distribution of the output

That's not out of distributíon, that's inside the distribution of the rollout. If you don't create rollouts for the game of Chess then it doesn't know how to play Chess no matter how smart it is at tasks you've created rollouts for. It's structurally stuck in its distribution.

Re: A recent experience with ChatGPT 5.5 Pro

#406

As a graduate student, this piece made me sad. I always believed that my work speaks for itself and transcends beyond my limited time on this cosmic experience. This notion of immortality was just a small intangible bonus I hoped for when I jumped into grad school. AI is making me feel less worthy.

"If you value intelligence above all other human qualities, you're gonna have a bad time." - Ilya Sutskever, 2023

Re: A recent experience with ChatGPT 5.5 Pro

#407

A very interesting comment from Baez, I'll just quote part of it. > Where does the value of thinking and having deep ideas come from? We need to think about this now. If it comes primarily from their scarcity – the fact that having certain ideas is hard – then indeed this value may drop precipitously when the manufacture of ideas can be automated. But if the value comes from the utility of the ideas – the benefit tha…

There are three species of mathematicians:

The first species is the pure problem solver. Tao is the poster child for this group. Their currency is interesting problems and solutions to those problems.

The second species is the pure theory builder. The poster child for this group is Conway. Their currency is theories and ideas rather than theorems, they are most interested in expanding the territory of mathematics and discovering new mathematical lands.

The third species is the applied mathematician. They see mathematics as a means to an end, they have some problem outside of mathematics and they want to use mathematics to solve it.

It seems like the first group (the problem solvers) are the most immediately threatened by AI, although so far AI is better at solving problems than finding new conjectures.

The second group (the theory builders) are more distantly threatened by AI, since thus far AI has shown limited ability to come up with novel and interesting mathematical ideas and nobody has any clue how to train an AI to do such a thing.

The third group stands to gain the most from AI. If an AI can answer your mathematical question then you can spend less time doing mathematics and more time on whatever it is outside of mathematics that you wanted to use mathematics to help solve.

Re: A recent experience with ChatGPT 5.5 Pro

#408

Earlier quoted context omitted.

> We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro Where's that? The stuff I've seen is from celebrities. Were those problems as hard as this one, or the ones that Tao posts about? Regardless.. what's the argument against more transparency here to just settle this kind of thing? > which is too conspiratorial of an explanation for my liking regardless OpenAI is…

> Where's that? https://archive.ph/2w4fi

> The duo had jump-started the AI-for-Erdős craze late last year by prompting a free version of ChatGPT with open problems chosen at random from the Erdős problems website. (An AI researcher subsequently gifted them each a ChatGPT Pro subscription to encourage their “vibe mathing.”)

Wonder who the AI researcher worked for? Is a "craze" something which a for-profit company would want to encourage? Maybe they'd think the publicity would help keep people talking about their company and product as we are now doing?

> “There was kind of a standard sequence of moves that everyone who worked on the problem previously started by doing,” Tao says. The LLM took an entirely different route, using a formula that was well known in related parts of math, but which no one had thought to apply to this type of question.

Yep, I do remember this now. Everyone was yelling that this was definitely a sign of ground-breaking and creative work, citing the expert. What the expert actually said suggests that the solution was available in training data! That also suggests the math in TFA is harder in comparison, answering my other question.

Pleasure as always HN, thanks for voting me to the bottom of the thread for this

Re: A recent experience with ChatGPT 5.5 Pro

#409
post #344

Earlier quoted context omitted.

If I was a frontier lab and I solved continual learning, as of today I would absolutely not release it - the society isn't ready for this; society isn't even ready for widespread diffusion of current publicly available frontier models. If however I was a frontier lab who solved continual learning and my competitor also solved and released it, I would release mine immediately, obviously. The point is, continual learni…

You're not a frontier lab, the shareholders own those. And if shareholders get a private briefing about an unprecedented breakthrough in continual learning, they would announce it from the rooftops to take credit for the progress ASAP and reap the rewards for their stock value. The only lab that I can exempt from this is DARPA.

Shareholders are not insiders. Public companies do secret projects all the time of which shareholders know absolutely nothing about and may never learn the details of them if they get cancelled.

Re: A recent experience with ChatGPT 5.5 Pro

#410
post #404

This jives with what I've experienced in the brief time I had access to 5.5 Pro. It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. The downside (not noted in the…

> This jives with what I've experienced Just as an fyi, the word you are looking for is jibes. Jive is something else entirely.

That ship sailed looooong ago.
Post reply on HN