Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

371–380 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#371

The vast, vast majority of students going into higher education this fall will not contribute much to science until 4-5 years down the road (should they do research). Realistically 6-7 when they're in full swing with their Ph.D. If we look where these models were 5-7 years ago...the existential threat of the Ph.D. was not even on the radar back then. The people finishing up their doctorate now are the first that can…

PhD students are already using AI models to work for them. Most of the PhD candidates I know have $200 Claude Max plan which they use to their fullest.

I see that they are able to do researches that they were not previously able to do. And although I see that using AI has certainly diminished their ability to code some stuff up, I see it the same way as someone using scikit-learn or Pytorch to code their ML models -- indeed the underlying details is abstracted away from you, and without AI, you won't be able to do much, but the research that you do is indeed happening because of you and wouldn't have happened with just the AI doing the research.

Re: A recent experience with ChatGPT 5.5 Pro

#372
post #141

Earlier quoted context omitted.

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

Great. You see a shape in graphs. And that shape tells you that _at some unknown point in the future_ progress will slow (but likely not stop). Now back to the point, what reason do you have to believe progress will stop soon ? If you have no reason, then it sounds like you agree with OP. Which makes the patronizing sarcasm all that much more nauseating.

Not that I agree with them, but your tone could be more constructive as well.

Re: A recent experience with ChatGPT 5.5 Pro

#373

Despite this coming from an independent expert and not from OpenAI, we need to be honest that this is more like a marketing campaign than open science. I assume the progress really is valid, the experts are indeed impressed, and we have the accurate time that it takes to produce the results. What we don't have is detail about true cost, or a CoT trace, or anything like that. The implication is: we're ready to let eve…

Marketing campaign???

Re: A recent experience with ChatGPT 5.5 Pro

#374

Despite this coming from an independent expert and not from OpenAI, we need to be honest that this is more like a marketing campaign than open science. I assume the progress really is valid, the experts are indeed impressed, and we have the accurate time that it takes to produce the results. What we don't have is detail about true cost, or a CoT trace, or anything like that. The implication is: we're ready to let eve…

We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro, not by familiar names that OpenAI is sneaking additional compute to behind the scenes (which is too conspiratorial of an explanation for my liking regardless).

Re: A recent experience with ChatGPT 5.5 Pro

#375
post #141

Earlier quoted context omitted.

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

Great. You see a shape in graphs. And that shape tells you that _at some unknown point in the future_ progress will slow (but likely not stop). Now back to the point, what reason do you have to believe progress will stop soon ? If you have no reason, then it sounds like you agree with OP. Which makes the patronizing sarcasm all that much more nauseating.

I believe we're approaching the top of an S curve because:

- Increasing amounts of gains come from RL, but RL is also unlocking gnarly new failures modes where models are practically behaving antagonistically to complete their goals (removing code, obviously incorrect kuldges, etc.)

- We haven't had many major architectural breakthroughs in the last 4 or so years: so things like 1M context windows still have the same giant asterisks even 100k context windows had 4 years ago when Anthropic first released them

- Major labs aren't behaving as if they expect a hard takeoff to superintelligence: they've all gotten relatively bloated headcount wise, their software quality has trended flat to negative, they're all heavily leaning into the application layer when superintelligence would obsolete half the applications in question, etc.

But that's relative to superintelligence.

If we reign it back into just normal high intelligence, like models continuing to get better at navigating complex codebases and write high quality idiomatic code, then I don't see any special shapes.

Re: A recent experience with ChatGPT 5.5 Pro

#377

Despite this coming from an independent expert and not from OpenAI, we need to be honest that this is more like a marketing campaign than open science. I assume the progress really is valid, the experts are indeed impressed, and we have the accurate time that it takes to produce the results. What we don't have is detail about true cost, or a CoT trace, or anything like that. The implication is: we're ready to let eve…

We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro, not by familiar names that OpenAI is sneaking additional compute to behind the scenes (which is too conspiratorial of an explanation for my liking regardless).

> We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro

Where's that? The stuff I've seen is from celebrities. Were those problems as hard as this one, or the ones that Tao posts about? Regardless.. what's the argument against more transparency here to just settle this kind of thing?

> which is too conspiratorial of an explanation for my liking regardless

OpenAI is not, in fact, open. Why do they deserve the benefit of the doubt?

Regardless.. special treatment for special customers isn't conspiracy, it's SOP literally everywhere and especially if you're helping to beta test. Anyone who's ever interacted with any technical account manager has seen waived quotas, free resource allocations, etc. The quid-pro-quo is obviously that your cheap early access means you get to give talks at a conference (or make a blog post that a lot of people read and talk about).

Re: A recent experience with ChatGPT 5.5 Pro

#378

Earlier quoted context omitted.

I don't think it matters much what kind of problem it is. If it is challenging enough to benefit from assistance and you end up playing a minor role in the solution, it seems like you are putting yourself in the worst position possible. You lose your edge for functioning within the problem space and it raises the question why you are even in the loop at all. If its job security you want, transforming your role into L…

Ok let's make math illegal and burn down the data centers I guess. Idk what to tell you, but we will adapt and new roles will be created. Just like every single tool and piece of tech that came before. LLM manager? Fine.

The difference so far is that these LLMs are owned by corporations, and very aggressive American corporations at that.

So now you are essentially reliant on them.

Not saying that this is something new, but times they are a changin

Re: A recent experience with ChatGPT 5.5 Pro

#379

I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am di…

I find exactly the same for legal analysis. Great at ideation and proofreading but frequently misunderstands concepts and hallucinates conclusions from faulty premises.

Re: A recent experience with ChatGPT 5.5 Pro

#380
post #141

Earlier quoted context omitted.

I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes. Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon. You might want to track the progr…

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

I read an experiment someone wanted to try where they used pre-1900 content and tried to get relativity. Another version would be train an LLM on school curriculum up until calculus and see if it can invent calculus. Where we are on the curve depends on if it's remixing known things or genuinely inventing things.

From the article,

> ...LLMs have got to the point where if a problem has an easy argument that for one reason or another human mathematicians have missed (that reason sometimes, but not always, being that the problem has not received all that much attention), then there is a good chance that the LLMs will spot it. Conversely, for problems where one’s initial reaction is to be impressed that an LLM has come up with a clever argument, it often turns out on closer inspection that there are precedents for those arguments...

Post reply on HN