Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

331–340 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#331

Earlier quoted context omitted.

@NotOscarWilde drop your email here, I will reach out and happy to get you a pro account for a few months so you can try 5.5 pro.(work at OAI)

While this sounds generous (and in some ways it is), it does not address the general point that GP is making. That is, the systematic disadvantage which large parts of humanity have w.r.t. to access to the tools. You could say they can't drive a Lambhorgini either, but that also doesn't solve the problem.

> While this sounds generous (and in some ways it is), it does not address the general point that GP is making. That is, the systematic disadvantage which large parts of humanity have w.r.t. to access to the tools. You could say they can't drive a Lambhorgini either, but that also doesn't solve the problem.

This was also the case historically, when being at certain universities, with better professors, better scope of works available at the library, etc, would necessarily provide systematic advantage.

This is the reality of progress. It is always unevely distrubuted.

I do think the open source side of model development is a substantial counter to the pessimism here.

Re: A recent experience with ChatGPT 5.5 Pro

#332
post #141

Earlier quoted context omitted.

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

It’s more of a guess if you don’t know about things like scaling laws and RL with verification. The onus of “we’re going to saturate” anytime soon is on that claim because every measurement points to that not being true.

But… RL doesn’t scale that well. It’s not the silver bullet you think it is.

Re: A recent experience with ChatGPT 5.5 Pro

#333
post #238

Earlier quoted context omitted.

In the sense that the incremental improvements in capabilities that we've been seeing in recent models seem to taking exponentially growing amounts of compute to achieve.

But they don't? Mythos is a 10T model. Opus is a 5T model. That's not an exponentially growing amount of compute but it is achieving exponential improvements (eg from Mozilla: https://blog.mozilla.org/en/privacy-security/ai-security-zer... )

where the heck did you get those parameter numbers from?

Re: A recent experience with ChatGPT 5.5 Pro

#334

Earlier quoted context omitted.

Yes, they can. Some people like to parrot "next token prediction", "LLMs can only interpolate", and other nonsense, but it is obviously not true for many reasons, in particular since we introduced RL. Humans do not have the monopoly on generating novel ideas, modern AI models using post training, RL etc can come to them in the same way we do, exploration. See also verifier's law [0]: "The ease of training AI to solve…

RL or no RL, AI cannot escape the distribution it's trained on. It's just that the labs will put so much into the distribution that we won't be able to tell the difference that easily, nor will it matter for most tasks. The reason AI does well on ARC-AGI-2 is because the labs created synthetic training data using similar puzzles.

Yes it can! That's the whole point of RL! it generates slightly out of distribution rollouts, and rewards good rollouts to change the distribution of the output

Re: A recent experience with ChatGPT 5.5 Pro

#335
post #35

Earlier quoted context omitted.

I mean in the same way getting Wolfram Alpha to solve a really hard/ugly differential equation I suppose

Mario Andretti could never have won a motor race without a car, yet we say he won the Indy500.

A motor race is defined by someone pilotting a motor car no?

We might not think we rightfully won an on foot race driving a car, yeah?

Re: A recent experience with ChatGPT 5.5 Pro

#336

As a TCS assistant professor from Eastern Europe, I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models. Paying for Pro from any of my current academic budgets is completely ouf of the field of reality here -- all budgets tend to have restricted uses and software payments fit into very few categories. Effectively, I'd have to ask for a brand new gr…

@NotOscarWilde drop your email here, I will reach out and happy to get you a pro account for a few months so you can try 5.5 pro.(work at OAI)

This doesn't solve the problem, though: having the ability to finish a field of study without paying a toll to a token provider.

Re: A recent experience with ChatGPT 5.5 Pro

#337
post #66
post #27

> Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would. This is a cultural choice. It makes sense that in the mathematics culture we currently hav…

I replied to a comment about AI in sports and I build on that. We praise car drivers despite most of the performance in their sport comes from the car. The driver makes the difference when two cars are close in performance. Brilliances or mistakes. Horse riders too. In the case of math, the human can lead the LLM on the right track, point it to a problem or to another one. So it deserves some praise. Then the team th…

Could you win an F1 race with the latest winning car against F1 drivers?

Re: A recent experience with ChatGPT 5.5 Pro

#338
post #273

Earlier quoted context omitted.

> When llms start contributing meaningfully to their own development, that would be a convincing indicator imo. This has been the case for awhile now already… https://kersai.com/the-48-hours-that-changed-ai-forever-clau...

And yet the world hasn’t changed all that much except people getting laid off in response to over-hiring prior to the diffusion of llm’s.

> over-hiring

For how long should you be allowed to use this excuse? It’s nearly 5 years since the peak of COVID hiring. What’s an acceptable limit - 10 years? Of course at that point you can just switch over to outsourcing and “stupid MBAs”, the other two of Reddit’s favorite scapegoats. I find a lot of the AI skepticism to be totally unfalsifiable.

Re: A recent experience with ChatGPT 5.5 Pro

#339
There have always been attempts at settling all mathematics by using mechanized approaches. Often by mathematicians who already had made an impact and then wanted an automated approach.

The Bourbaki group was one of the first who attempted a mechanized approach (using pen and paper still of course) to set theory and were literally accused of wanting to end all mathematics. The approach was largely ignored in practice.

Gowers and a handful of others who work on computerized approaches also seem to want to end human mathematics and have sharecropper mathematics for a monthly tithe. So far they are largely ignored in practice.

Re: A recent experience with ChatGPT 5.5 Pro

#340
post #141

Earlier quoted context omitted.

I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes. Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon. You might want to track the progr…

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

you can tell where on the sigmoid we're currently sitting? frontier lab folks can't - chapeau bas good sir
Post reply on HN