Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

361–370 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#361

Earlier quoted context omitted.

Great. You see a shape in graphs. And that shape tells you that _at some unknown point in the future_ progress will slow (but likely not stop). Now back to the point, what reason do you have to believe progress will stop soon ? If you have no reason, then it sounds like you agree with OP. Which makes the patronizing sarcasm all that much more nauseating.

Nausea aside, what evidence does anyone have that “super intelligence” of the sort your argument alludes to is even possible? Because that’s what we’re really talking about; greater than human intelligence on this sort of academic task. For example; When llms start contributing meaningfully to their own development, that would be a convincing indicator imo.

Well, a decent GPU runs on 20x the wattage of a human brain. That's evidence humans are constrained in ways artificial intelligences will not be.

Re: A recent experience with ChatGPT 5.5 Pro

#362
post #354

Earlier quoted context omitted.

> over-hiring For how long should you be allowed to use this excuse? It’s nearly 5 years since the peak of COVID hiring. What’s an acceptable limit - 10 years? Of course at that point you can just switch over to outsourcing and “stupid MBAs”, the other two of Reddit’s favorite scapegoats. I find a lot of the AI skepticism to be totally unfalsifiable.

And the same can be said for AI exuberance. Yes, LLMs are a great technology. Yes, we will probably all use them all the time in 20 years. No, we don't know how we will use them (to generate cat memes or to cure cancer) in 20 years time. Especially for software developers it looks increasingly that after huge turmoil it's likely we will need +/- the same number of developers in the world.

> Especially for software developers it looks increasingly that after huge turmoil it's likely we will need +/- the same number of developers in the world.

what exactly are you basing this opinion on? All I am seeing personally across multiple projects I am working on and other friends at other places is that downsizing is either begun or is planned (to exclude from here all the “public” layoffs we see on the news). Given how most business operate in the USA I think most of “AI strategies” are “we can do same with -40% staff” vs. “we can do XX% more work with same staff.”

Re: A recent experience with ChatGPT 5.5 Pro

#363
post #141

Earlier quoted context omitted.

I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes. Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon. You might want to track the progr…

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

Assuming it’ll stop soon is to wager that we’re at a very specific point on the curve.

If it’s anyone’s guess then we’re much more likely to be left of that, unless you argue we’re already on the flat side.

Re: A recent experience with ChatGPT 5.5 Pro

#364

Earlier quoted context omitted.

I don’t think it’s just mathematics. We don’t hear enough about this, but if I think back to my undergraduate years, which were less than 10 years ago, every homework assignment and every take-home exam I had would be trivial for LLMs to solve at this point I wonder what is actually happening on the ground.

Well... here's something from "boots on the ground": I teach a bachelor's degree where programming is a smallish facet of a curriculum. My course is the last of a series of 3 courses which progressively introduce more concepts and try make practical implementations more feasible. I've been able to grade the course purely based on returns to take-home exercises, some of which are complex, some trivial. When ChatGPT (&…

I had this exact problem and came to the same conclusion. But for the exam, give them code and ask what it does, or give broken code and ask where the error is. Waaay less marking. I only asked for 3 small functions written by hand and that was still 90% of the effort to mark. But the marks felt valid in the end, so the process seemed to work.

Re: A recent experience with ChatGPT 5.5 Pro

#365

Earlier quoted context omitted.

Anthropomorphism is a subtle marketing tool used by these big AI companies, who are financially incentivized to push the myth of AGI and want everyone to believe they're right on the cusp of achieving it. It's good to be pedantic in this case, we shouldn't anthropomorphize these tools.

This is just a “hurr durr AI companies evil” argument without substance. It’s the people that are the problem, nobody told the grandparent to use “mentoring” as a word, and my argument is that it’s a complete overreaction to classify them as anthropomorphizing AIs, and I’d argue default to that argument would be an insult to them, and it’s super pedantic.

> This is just a “hurr durr AI companies evil” argument without substance.

If you say so bud.

> nobody told the grandparent to use “mentoring” as a word

Nobody told people to say "Google it" either; nobody told us to use the word "Kleenex" when we mean tissue; nobody told us to use the word "Chapstick" when we mean lip balm. Nobody told British people to say "Hoover" when they mean vacuum, or "Sellotape" when they mean transparent tape.

This is literally how soft influence works, it's how brands "colonize" language. A professor using the anthropomorphized word "mentoring" when talking about a machine, as if it's a student that can learn and develop relationships, is this same soft influence at work. The AI companies' websites are all riddled with cognitive language, their chat bots all use conversational UI like you're talking to a person, the bots answer with "we," "me," and "I." They created an environment that made anthropomorphized language feel natural, which only helps their marketing goals.

Go ahead and call it pedantry all you want, but that's the whole point. The problem is epistemic.

Re: A recent experience with ChatGPT 5.5 Pro

#366

Earlier quoted context omitted.

Hmm, I don’t know, maybe the fact that 4.6, 4.7, 5.3, 5.4, 5.5, 3.0, 3.1 are all marginal improvements?

I think people's opinion of "marginal improvement" is based on their relative ability. A 2000 elo chess player is going to think the jump from 500 to 1000 is marginal. They're both floundering around not doing anything resembling common sense. A 1000 elo chess player is going to find the jump from 2000 to 2500 marginal. They're both playing far better moves for incomprehensible reasons, and the only reason you know t…

2024-2025 was filled with huge improvements. 2025-2026 has not been, outside of open source.

The idea that we’re at the point where it’s superseded our ability to tell just makes no sense. I’ll be happy if we can get to a point where I don’t have to tell Claude not to tail every bash command or make a job that writes throughout instead of once at the end. I’ll be happy if “continue this interaction naturally, you are taking over from an independent subagent” works.

But I’m not holding my breath. It’s still really cool that any of this stuff is possible.

Re: A recent experience with ChatGPT 5.5 Pro

#367

Earlier quoted context omitted.

Are you a cutting edge research scientist or something? Everyone I know works in the same domain every day. The problems are the same. People aren't solving brand new problems to humanity every day. We make budgets and look at ticket counts. Roll out patches. Replace hardware. Upgrade software packages. Make a new dashboard to track a project. I guess if every day is a completely novel thing for you, ok. I feel like…

I don't think it matters much what kind of problem it is. If it is challenging enough to benefit from assistance and you end up playing a minor role in the solution, it seems like you are putting yourself in the worst position possible. You lose your edge for functioning within the problem space and it raises the question why you are even in the loop at all. If its job security you want, transforming your role into L…

Ok let's make math illegal and burn down the data centers I guess. Idk what to tell you, but we will adapt and new roles will be created. Just like every single tool and piece of tech that came before. LLM manager? Fine.

Re: A recent experience with ChatGPT 5.5 Pro

#368

Earlier quoted context omitted.

Hmm, I don’t know, maybe the fact that 4.6, 4.7, 5.3, 5.4, 5.5, 3.0, 3.1 are all marginal improvements?

Equally marginal?

No, the anthropic releases have felt marginally negative

Re: A recent experience with ChatGPT 5.5 Pro

#369
Despite this coming from an independent expert and not from OpenAI, we need to be honest that this is more like a marketing campaign than open science. I assume the progress really is valid, the experts are indeed impressed, and we have the accurate time that it takes to produce the results. What we don't have is detail about true cost, or a CoT trace, or anything like that.

The implication is: we're ready to let everyone go wild with this very soon. Ok, go wild with what exactly? How do we know that influential VIP users who might make very friendly blog posts aren't getting allocated exclusive access to a billion dollars worth of hardware when they ask questions? I mean literally giving a certain group of people temporary privileged access to like 90% of all available compute would be a completely reasonable business decision for OpenAI.

Would a reveal like that change how we think about the result? What if half that amount of cash/compute could enable some completely non-AI approach of numerical brute forcing that settles the question even if it didn't write the paper?

My other question is always whether the latest is purely using giant models or if we're now deeply into harnesses that use MCTS and such. Understandable to keep that a trade secret I guess. But IMHO we should at least get the CoT trace as a proxy for true cost, or else maybe we're just getting played to do the hype for corporate.

Re: A recent experience with ChatGPT 5.5 Pro

#370

Earlier quoted context omitted.

This is just grossly misinformed. OAI and Anthropic both require KYC for models of similar intelligence. They both do account bans if the classifiers fire wrong. You simply hear about it less with OAI because Codex has fewer prosumers.

Can you name any instance of OpenAI being as trigger-happy with bans as Anthropic has been in the past few months? Codex may have fewer prosumers, but they've added a lot in that time.

There will obviously be more reports about Anthropic because hating on Anthropic has been trendy lately, it has more users, and gullible people (including most HN'ers) fall victim to OAI's guerrilla marketing.

In terms of bans and KYC they are not meaningfully different.

Post reply on HN