Live data from Hacker News

A recent experience with ChatGPT 5.5 Pro

gowers.wordpress.com

551–558 of 558 posts

Re: A recent experience with ChatGPT 5.5 Pro

#551

Earlier quoted context omitted.

The cost factors on the new models compared to the old models.

Cost for a specific level of performance decreases 10x per year, this has been a pretty consistent property for awhile now.

I guess within the domain of AI, a pertinent question would be: "do I want to use anything but the best?" The errors older models give being directly analogous to being stupider in my eyes.

Re: A recent experience with ChatGPT 5.5 Pro

#552

Earlier quoted context omitted.

Are you using this agent hive for any repeatable tasks? What you described, superficially, seems like a one off. Genuinely curious.

I think it depends on what you mean by repeatable tasks. I reuse the critical handoff agents quite a lot since they are basically just set up to help spot bias and errors. I kind of reuse the top agent. I have a few "core" configurations that I can add to. So one will know our network, one will know our data architecture and so on, to keep them a little more focused. So for this specific agent that I described, I'll…

In my previous jobby job I needed to pull CSVs out of Tableau, then from an ancient monolithic PHP admin and other sources, then manually merge them, reformat them in G Sheets, pivot this and that and send the report to my supervisor. Initially took 2 hours then down to one, but still senseless busy work. It was the “fault” of the incumbent IT, but if I could turn that hour into a minute… I wouldn’t get a raise, but I’d have more time for something else or nothing. I feel like this is still a scenario for countless many and perhaps the valley of the low hanging fruit. That’s where my question was coming from.

Your firm seems to operate on a higher plane, jealous :)

Re: A recent experience with ChatGPT 5.5 Pro

#553

Earlier quoted context omitted.

Cost for a specific level of performance decreases 10x per year, this has been a pretty consistent property for awhile now.

I guess within the domain of AI, a pertinent question would be: "do I want to use anything but the best?" The errors older models give being directly analogous to being stupider in my eyes.

Depends — many tasks in various pipelines have a reasonable Pareto frontier and diminishing returns after a certain level of performance. You may just have a high budget constraint (say like YouTube computing ASR subtitles; they are not going to be using the best ASR models because it’s expensive). If it’s myself, with a coding agent, I’m going to get the best thing I can afford.

Re: A recent experience with ChatGPT 5.5 Pro

#554

All this sounds to me like mathematicians spooking themselves with stories of how ChatGPT solved a problem, when it's mathematicians solving a problem using ChatGPT as a tool. E.g. from the twitter thread by Timothy Gowers: >> All I did was say things like, "Yes, it would be great if you could explore that idea and see whether you can get it to work," or "Could you rewrite that argument as a LaTeX file in the style o…

>If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening.

I mean that has happened so yeah ?

https://www.scientificamerican.com/article/amateur-armed-wit...

Actual GPT transcript. Zero such input https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...

And maybe the other guy wasn't the most polite about it but his point is very valid. Replace chatgpt with a human in both of these stories and nobody would say that timothy 'took the horse and made it drink'. The 'Horse' would be the first and likely only Author so this just sounds like denial.

That there are multiple of these stories in the last few months by the latest set of models (there are even more than these 2) should provoke this sort of consideration and discussion.

Re: A recent experience with ChatGPT 5.5 Pro

#555

All this sounds to me like mathematicians spooking themselves with stories of how ChatGPT solved a problem, when it's mathematicians solving a problem using ChatGPT as a tool. E.g. from the twitter thread by Timothy Gowers: >> All I did was say things like, "Yes, it would be great if you could explore that idea and see whether you can get it to work," or "Could you rewrite that argument as a LaTeX file in the style o…

>If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening. I mean that has happened so yeah ? https://www.scientificamerican.com/article/amateur-armed-wit... Actual GPT transcript. Zero such input https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba... And maybe the other guy wasn't the most polite about it but his point…

These are different cases, yes? The person in the SA article you link is described as an "amateur", but Timothy Gowers is not an amateur and he is much more capable of guiding an LLM with domain expertise than an amateur.

Then there's the kind of problem we're talking about. The "amateur" in the SA article solved one of Erdős problems and Gowers himself seems to think that, on its own, is not a cause for concern. He distinguishes his own result from that kind of earlier result at the start of his article:

>> The background is that, as has been widely reported, LLMs are now capable of solving research-level problems, and have managed to solve several of the Erdős problems listed on Thomas Bloom’s wonderful website. Initially it was possible to laugh this off: many of the “solutions” consisted in the LLM noticing that the problem had an answer sitting there in the literature already, or could be very easily deduced from known results.

So we have an "amateur" who "vibe-solved" an Erdős problem, on one hand, which may or may not already had a solutiuon lurking in the wings on the one hand; and an expert who solved a harder problem by interactive use rather than vibe-solving, on the other hand. There's no reason to believe that we can "Replace chatgpt with a human in both of these stories" as you say.

And btw there's scholarship that indicates vibe-solving is not yet ready to replace mathematicians like Timothy Gowers:

First Proof

To assess the ability of current AI systems to correctly answer research-level mathematics questions, we share a set of ten math questions which have arisen naturally in the research process of the authors. The questions had not been shared publicly until now; the answers are known to the authors of the questions but will remain encrypted for a short time.

https://arxiv.org/abs/2602.05192

See Appendix A for initial results.

Re: A recent experience with ChatGPT 5.5 Pro

#556

Earlier quoted context omitted.

Watching a teenager approach their homework, instead of struggling to answer questions they don't know, they ask Gemini. Unfortunately, I think the mental struggle to approach an answer is where much of the learning is. They also miss out on the reward for persistence of seeing things fall together. It is troubling. It suggests a plateauing of human understanding.

It absolutely is where the learning is, that's pretty well established brain science.

It's a struggle to map a good approach - LLM-based tools are a boon in some respects, like having a personal tutor. But in others are fundamentally opposed to the process of education.

Like, I asked ChatGPT to make me some problems, it did, then I got to check my answers. In the past I'd have had a textbook for that; but schools stopped giving those out decades ago.

Re: A recent experience with ChatGPT 5.5 Pro

#557

Earlier quoted context omitted.

>If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening. I mean that has happened so yeah ? https://www.scientificamerican.com/article/amateur-armed-wit... Actual GPT transcript. Zero such input https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba... And maybe the other guy wasn't the most polite about it but his point…

These are different cases, yes? The person in the SA article you link is described as an "amateur", but Timothy Gowers is not an amateur and he is much more capable of guiding an LLM with domain expertise than an amateur. Then there's the kind of problem we're talking about. The "amateur" in the SA article solved one of Erdős problems and Gowers himself seems to think that, on its own, is not a cause for concern. He…

Yes these are different instances.

My first point is that I think you are overating 'interactive use' a bit here. Like Timothy already explains in the article, Were it a human he 'guided' in a similar way, he would not get credit for those achievements by any stretch of the imagination. And I think that's an important part of realizing why these sort of people are beginning to discuss these things.

Second. I didn't say anything about models being ready to replace mathematics wholesale. But should people really wait until that happens before discussing it? I know it's human nature to wait until the problem or situation is upon you but I don't think that would be prudent or wise. And even just for the sake of curiosity, it would be boring.

I think the matter of fact here is that in the last few months with the last few models, capabilities in this area have jumped to a very meaningful degree. It would be stranger if no one was talking about it.

Re: A recent experience with ChatGPT 5.5 Pro

#558
post #482

Earlier quoted context omitted.

> It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. I swear that people have said the same thing with effectively every new model that came out in the last six mo…

They did. The scam continues.

The conspiracy theory mindset, is there anything it can't explain away?
Post reply on HN