"After 16 minutes and 41 seconds, it came back" ... "further 47 minutes and 39 seconds" ... "After 13 minutes and 33 seconds" ... "After 9 minutes and 12 seconds" ... "After 31 minutes and 40 seconds" ... plus other computations Anyone spotting the issue here? What did that really cost? I am not against compute being used for scientific or other important problems. We did that before LLMs. However, the major LLM gate…
A recent experience with ChatGPT 5.5 Pro
81–90 of 558 posts
Re: A recent experience with ChatGPT 5.5 Pro
#82Earlier quoted context omitted.
@NotOscarWilde drop your email here, I will reach out and happy to get you a pro account for a few months so you can try 5.5 pro.(work at OAI)
This requires a major "dox" of myself, but I am really grateful for the offer, so these are my academic contacts: https://pastebin.com/hNYrCjhL I probably will erase the contents in a few days. Even if you just drop an email and it doesn't work out, I appreciate this gesture so much. Thank you.
Re: A recent experience with ChatGPT 5.5 Pro
#83I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am di…
please, sign up for a paid plan of either chatgpt or claude. gemini is while close, still noticeably behind you deserve opinions shaped by interactions with the best tools that are out there.
Re: A recent experience with ChatGPT 5.5 Pro
#84I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am di…
Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon.
You might want to track the progress of these models on the CritPt benchmark, which is built on *unpublished, research-level* physics problems:
Frontier models are still nowhere near solving it, but progress has been rapid.
* o3 (high) * GPT 5.4 (xhigh), 23.4%
* GPT-5.5 (xhigh), 27.1%
* GPT-5.5 Pro (xhigh) 30.6%.
Re: A recent experience with ChatGPT 5.5 Pro
#85I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked. However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am di…
I put my stuff through several sota models and round robin them in adversarial collaboration and they are all useful even though, fundamentally, they don’t “understand” anything. But they are super useful delegates as long as deciding on the problem and approach and solution all sits safely in your head so you can challenge them and steer them.
So I know the article is about one particular new model acing something and each vendor wants these stories to position their model as now good enough to replace humans and all other models, but working somewhere where I am lucky enough to be able to use all the sota models all the time, I can say that all keep making obvious mistakes and using all adversarially is way better than trusting just one.
I look forward to the day one a small open model that we can run ourselves outperforms the sum of all today’s models. That’s when enough is enough and we can let things plateau.
Re: A recent experience with ChatGPT 5.5 Pro
#86"After 16 minutes and 41 seconds, it came back" ... "further 47 minutes and 39 seconds" ... "After 13 minutes and 33 seconds" ... "After 9 minutes and 12 seconds" ... "After 31 minutes and 40 seconds" ... plus other computations Anyone spotting the issue here? What did that really cost? I am not against compute being used for scientific or other important problems. We did that before LLMs. However, the major LLM gate…
Still not as bad for the environment as animal agriculture, and animal agriculture is absolutely not necessary and only causes harm and suffering for taste pleasure. At least with LLMs we get many positive advancements from them. I don't see these sorts of comments every time someone posts a burger review.
Re: A recent experience with ChatGPT 5.5 Pro
#87As a TCS assistant professor from Eastern Europe, I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models. Paying for Pro from any of my current academic budgets is completely ouf of the field of reality here -- all budgets tend to have restricted uses and software payments fit into very few categories. Effectively, I'd have to ask for a brand new gr…
Re: A recent experience with ChatGPT 5.5 Pro
#88Re: A recent experience with ChatGPT 5.5 Pro
#89Link to source blog post: https://gowers.wordpress.com/2026/05/08/a-recent-experience-...
Re: A recent experience with ChatGPT 5.5 Pro
#90As a TCS assistant professor from Eastern Europe, I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models. Paying for Pro from any of my current academic budgets is completely ouf of the field of reality here -- all budgets tend to have restricted uses and software payments fit into very few categories. Effectively, I'd have to ask for a brand new gr…
@NotOscarWilde drop your email here, I will reach out and happy to get you a pro account for a few months so you can try 5.5 pro.(work at OAI)