Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

341–350 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#341
post #141

Earlier quoted context omitted.

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

Great. You see a shape in graphs. And that shape tells you that _at some unknown point in the future_ progress will slow (but likely not stop). Now back to the point, what reason do you have to believe progress will stop soon ? If you have no reason, then it sounds like you agree with OP. Which makes the patronizing sarcasm all that much more nauseating.

Hmm, I don’t know, maybe the fact that 4.6, 4.7, 5.3, 5.4, 5.5, 3.0, 3.1 are all marginal improvements?

Re: A recent experience with ChatGPT 5.5 Pro

#342
post #97

Earlier quoted context omitted.

Gemini feels deep and philosophical. Especially for product management. Tell him you're a product manager and we're a team of two. But regular reminder - All LLMs can be wrong all the time. I only work with LLMs in domains I'm expert in OR I have other sources to verify their output with utmost certainty.

> I only work with LLMs in domains I'm expert in This. Should become a general rule for any non-trivial use of LLM in a professionel setting.

LLMs can also be really good in fields where you are not an expert. You just need to be very aware of your limitations, and start parallel conversation so one agent fact checks the other.

Re: A recent experience with ChatGPT 5.5 Pro

#343
A very interesting comment from Baez, I'll just quote part of it.

> Where does the value of thinking and having deep ideas come from? We need to think about this now. If it comes primarily from their scarcity – the fact that having certain ideas is hard – then indeed this value may drop precipitously when the manufacture of ideas can be automated. But if the value comes from the utility of the ideas – the benefit that the idea brings – then the story changes: perhaps creating more good ideas is actually better, not worse. Here I’m using “utility” in a broad sense, not just in the sense of what people often call applied mathematics.

> In other words, mathematicians may need to adjust to a transformation from a scarcity economy to an abundance economy.

https://gowers.wordpress.com/2026/05/08/a-recent-experience-...

Re: A recent experience with ChatGPT 5.5 Pro

#344

Earlier quoted context omitted.

Using the word “Mentoring” is anthropomorphic and subconsciously makes you think it will learn. It does not, and it is for the human brain a formidable task to remember that something as smart as an LLM does not learn. I keep catching myself making the same mistake. It’s also because it is so annoying to have to manage the memory of the LLM with custom prompts/instructions manually. I have not yet played with the lon…

Current LLM architecture doesn't learn - and you're right this is a huge piece that normal folks fail to understand, since in many ways, it's the opposite of what years of AI research has been trying to create. However, I think it's important to remember that LLMs are embedded in larger systems, and those larger systems do learn.

If I was a frontier lab and I solved continual learning, as of today I would absolutely not release it - the society isn't ready for this; society isn't even ready for widespread diffusion of current publicly available frontier models.

If however I was a frontier lab who solved continual learning and my competitor also solved and released it, I would release mine immediately, obviously.

The point is, continual learning might be solved already, we just don't know and those who might know would rather keep their mouths shut. It isn't my base case (financial situation of frontier labs is such that they'd probably release immediately as long as they have inference compute to serve this revolutionary capability), but it isn't impossible.

Re: A recent experience with ChatGPT 5.5 Pro

#345
post #340
post #141

Earlier quoted context omitted.

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

you can tell where on the sigmoid we're currently sitting? frontier lab folks can't - chapeau bas good sir

> frontier lab folks can't

Do you have a source for this that isn't marketing spiel? There's a fiscal incentive to lie about scaling research.

Re: A recent experience with ChatGPT 5.5 Pro

#346
Like coding, if you get inspired by AI for a novel idea, and can reproduce the same result independently (could code the same thing by hand) or at least understand and check every single argument (self review your code, test on your machine) and get it peer reviewed (code review, but with a real human) then I don’t see why the industry accepts the latest iteration of ChatGPT being 99% written by codex, but rejects a valid math result inspired by it.

Re: A recent experience with ChatGPT 5.5 Pro

#348
post #234

Earlier quoted context omitted.

Yes, they can. Some people like to parrot "next token prediction", "LLMs can only interpolate", and other nonsense, but it is obviously not true for many reasons, in particular since we introduced RL. Humans do not have the monopoly on generating novel ideas, modern AI models using post training, RL etc can come to them in the same way we do, exploration. See also verifier's law [0]: "The ease of training AI to solve…

Reinforcement learning for "reasoning" perturbs the model to generate completions in a particular chain of thought / alternative selection structure. It's three next token predictors in a trench coat.

When these things start solving many more long standing problems, and start introducing more novel problems, will people finally admit that the "next token predictor" is not the gotcha they think it is?

Re: A recent experience with ChatGPT 5.5 Pro

#349

I think the biggest advantage of ChatGPT compared to Claude is that there are fewer things outside the model itself, such as KYC, account bans, etc.

This is just grossly misinformed. OAI and Anthropic both require KYC for models of similar intelligence. They both do account bans if the classifiers fire wrong. You simply hear about it less with OAI because Codex has fewer prosumers.

Can you name any instance of OpenAI being as trigger-happy with bans as Anthropic has been in the past few months? Codex may have fewer prosumers, but they've added a lot in that time.

Re: A recent experience with ChatGPT 5.5 Pro

#350

A very interesting comment from Baez, I'll just quote part of it. > Where does the value of thinking and having deep ideas come from? We need to think about this now. If it comes primarily from their scarcity – the fact that having certain ideas is hard – then indeed this value may drop precipitously when the manufacture of ideas can be automated. But if the value comes from the utility of the ideas – the benefit tha…

I note that it is always the same online pundits (even if they are distinguished academics) who push anything new.

Meanwhile Wiles and Perelman stayed offline and solved real problems.

Post reply on HN