Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

211–220 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#211
post #141

Earlier quoted context omitted.

I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes. Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon. You might want to track the progr…

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

There are advancements that do not follow s curves - consider for instance total data transmitted over all networks, or financial derivatives volumes.

I think a better question for AI is “is it more like a network effect, liquidity effect, or a biological/physical effect”?

Re: A recent experience with ChatGPT 5.5 Pro

#212

Earlier quoted context omitted.

Exactly, if I generate a large chunk software, I'm going to have expectations about what it will do, how it will do it, etc. You don't just accept the statement that "it's done" for fact but you start looking for evidence. A scientific approach here is to look to falsify the statement. You start asking questions, running tests, experiments, etc. to prove the notion that it is done wrong. And at some point you run out…

> Mostly I just nudge it along. "Did you think about X? What about Y? Let's test Z" Exactly - you need to constantly have your sceptics glasses on and you need to be exacting in terms of the structure you want things to follow. Having and enforcing "taste" is important and you need to be willing to spend time on that phase because the quality of the payoff entirely depends on it. I recently planned for a major refact…

So you have to know the answer and also be an expert in the problem domain?

Re: A recent experience with ChatGPT 5.5 Pro

#213

[flagged]

I don’t love the tone here, but I do think you get at a key question in mathematical philosophy.

Mathematicians have engaged, vigorously, on this very philosophical question for centuries - is math discovered truth, or is it more akin to building an edifice where you first define the materials, then the structure, and see where it leads?

There are lots of strong feelings on both sides. For instance: “God created the integers, the rest is the creation of man” — Kronecker, 19th century sums up one particular perspective.

To me, it’s probably a mix of both - some fantastic results in imaginary numbers show up as describing key electromagnetic effects many decades after they were first ‘discovered’ by theoretical mathematicians.

NB: My original comment led with a pejorative, which was rightly flagged.

Re: A recent experience with ChatGPT 5.5 Pro

#214

Earlier quoted context omitted.

Using the word “Mentoring” is anthropomorphic and subconsciously makes you think it will learn. It does not, and it is for the human brain a formidable task to remember that something as smart as an LLM does not learn. I keep catching myself making the same mistake. It’s also because it is so annoying to have to manage the memory of the LLM with custom prompts/instructions manually. I have not yet played with the lon…

Current LLM architecture doesn't learn - and you're right this is a huge piece that normal folks fail to understand, since in many ways, it's the opposite of what years of AI research has been trying to create. However, I think it's important to remember that LLMs are embedded in larger systems, and those larger systems do learn.

exactly like you said - the harness might learn.

we do also have training on synthetic data. it might compound.

Re: A recent experience with ChatGPT 5.5 Pro

#216
post #27

> Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would. This is a cultural choice. It makes sense that in the mathematics culture we currently hav…

I would. Even if someone found a prompt or even automated the conversation and just searched all open math problems I still would. If they produced a useful result without harm to anyone, that's a valuable human activity that should be rewarded just as well as we reward the other mathematicians, which I imagine is quite a lot, given all the billionaire mathematicians...

Re: A recent experience with ChatGPT 5.5 Pro

#219

I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am di…

please, sign up for a paid plan of either chatgpt or claude. gemini is while close, still noticeably behind you deserve opinions shaped by interactions with the best tools that are out there.

Agreed, Gemini is clearly a capable model, but the tool use is lagging behind the other two. Ironically it regularly gets things wrong (ie. the current version of some software) because of an unwillingness to use web search.

Re: A recent experience with ChatGPT 5.5 Pro

#220
post #141

Earlier quoted context omitted.

> there's no reason to believe the progress of LLMs [...] will stop anytime soon Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

There are advancements that do not follow s curves - consider for instance total data transmitted over all networks, or financial derivatives volumes. I think a better question for AI is “is it more like a network effect, liquidity effect, or a biological/physical effect”?

Those are measuring the utility of a technological advancement by looking at usage, not the pace of advancement of said technology.
Post reply on HN