Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

241–250 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#241
post #141

Earlier quoted context omitted.

I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes. Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon. You might want to track the progr…

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

Software and hardware have no limits. Theoretically would could bozons for computations and have the same amount of computation available on one cm3 of the current total computation in the entire world. Same with software. Never there was a stop on new algorithms. With LLMs there are so many parts that will get better and are not very far fetched.

Re: A recent experience with ChatGPT 5.5 Pro

#242
post #237

Earlier quoted context omitted.

If higher bandwidth networking consisted primarily running more and more ethernet lines in parallel, you would most certainly agree that "networking has stagnated". "Reasoning" and now "Agentic" AI systems are not some fundamental improvement on LLMs, they're just running roughly the same prior-gen LLMS, multiple times. Hence the conclusion that LLM improvement has slowed down, if not stagnated entirely, and that we…

From TFA: “ChatGPT came up with an idea which is original and clever. It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour to find and prove”

You misunderstand. I'm not saying that Reasoning/Agentic systems aren't better.

I'm saying they're not an advancement in the tech in the way GPT 1 through 3 were. They're a different kind of improvement.

And as such the rate improvement cannot just be extrapolated into the future.

Re: A recent experience with ChatGPT 5.5 Pro

#244

I think the biggest advantage of ChatGPT compared to Claude is that there are fewer things outside the model itself, such as KYC, account bans, etc.

This is just grossly misinformed.

OAI and Anthropic both require KYC for models of similar intelligence. They both do account bans if the classifiers fire wrong. You simply hear about it less with OAI because Codex has fewer prosumers.

Re: A recent experience with ChatGPT 5.5 Pro

#245

Earlier quoted context omitted.

I agree and put it this way: LLMs sound so convincing presenting you the work it does rose colored and promising to give you more if you keep going. There is a 50/50 chance that it turns out to be right or letting you jump of the cliff. Only the trip stays the same beautiful 5 star plus travel. Also, spotting an error and telling LLM makes it in most cases worse, because the LLM wants to please you and goes on to apo…

Reusing the same prompt several times is something I've started doing too. The contrast is often illuminating. In one case, it made a thoroughly convincing argument that an approach was justified. The second time it made exactly the opposite argument, which was equally compelling. I now see LLMs as persuasion machines.

Before AI happened I watched youtube. Occasionally I encountered there very convincing arguments. Same person often made very convincing arguments on many subjects.

But noticed that the closer the domain they were talking about was to my area of competence the less convincing their arguments were. There were more holes, errors and wrong conclusions.

I recalibrated my bs meter thanks to that.

Since AI came I successfully used this strategy of being extremely cautious towards convincing arguments to not become mislead by AI.

However this year I'm working with AI more in the domain of software development. Where I can see the competence. And I see the competence. This had opposite effect on me. I tend to trust AI outside my domain of expertise much more after I saw what can it do in software.

One caveat though is that there are a lot of areas of human culture where there's very little actual knowledge, but a lot of opinions, like politics, economy, diet, business, health. I still don't trust AI in those domains. But then again, I don't trust humans there either.

For me basically AI achieved the threshold of useful reliability for any domain that humans are reliable at.

I don't really care about sycophancy. I might have a slight advantage that I don't talk to AI in my native language. So its responses don't have a direct line to my emotions.

Re: A recent experience with ChatGPT 5.5 Pro

#246
post #6

It's a very long post with a mix of technical (math) and philosophical sections. Here are the most striking points to reflect upon IMHO. > It seems to me that training beginning PhD students to do research [...] has just got harder, since one obvious way to help somebody get started is to give them a problem that looks as though it might be a relatively gentle one. If LLMs are at the point where they can solve “gentl…

> Would we regard that as a major achievement of the mathematician? I don’t think we would.

1. Does it matter, really? 2. Is it very different from previous computer-aided proofs, philosophically?

Re: A recent experience with ChatGPT 5.5 Pro

#247

I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am di…

LLM’s are the most powerful tool invented to search across a huge information space in response to human input.

That’s all they are. They don’t ‘know’ anything intrinsically and do know ‘know’ what reasoning even is.

Re: A recent experience with ChatGPT 5.5 Pro

#248

Earlier quoted context omitted.

Insane that we have a system capable of making innovative math proofs and people dismiss it as unimpressive

The creation of the system is deeply impressive, so are compilers but I don't raise a toast to it each time I build my code. Like generated art, people aren't going appreciate it on the same level.

Wow you consider this on the same level of impressiveness as a compiler?

Re: A recent experience with ChatGPT 5.5 Pro

#249
post #7

>So if your aim in doing mathematics is to achieve some kind of immortality, so to speak, then you should understand that that won’t necessarily be possible for much longer — not just for you, but for anybody. This made me a little sad

I watched the movie '21' (2008) for free on YouTube yesterday.

The opening of the movie features the MIT campus full of students navigating its grounds and all the promise and status that higher education brings. [0]

Gave me the same sense of sadness realizing how much will fall to AI.

[0] - https://youtu.be/0lsUsWdkk0Y?si=TJl7f_b1RcWcDqF8&t=278

Re: A recent experience with ChatGPT 5.5 Pro

#250
post #237

Earlier quoted context omitted.

From TFA: “ChatGPT came up with an idea which is original and clever. It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour to find and prove”

You misunderstand. I'm not saying that Reasoning/Agentic systems aren't better . I'm saying they're not an advancement in the tech in the way GPT 1 through 3 were. They're a different kind of improvement. And as such the rate improvement cannot just be extrapolated into the future .

GPT1 through GPT3 advancement were exactly like using more Ethernet cables in parallel.

All interesting conceptual breakthroughs came after GPT3: RL and reasoning being the main ones.

Post reply on HN