Earlier quoted context omitted.
My apologies. I’m out of the edit window, so I can’t remove it, but I tried!
Appreciated! I've re-opened it for editing if you want to. For us the main point is just to fix things going forward!
A recent experience with ChatGPT 5.5 Pro
511–520 of 558 posts
Re: A recent experience with ChatGPT 5.5 Pro
#512As a TCS assistant professor from Eastern Europe, I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models. Paying for Pro from any of my current academic budgets is completely ouf of the field of reality here -- all budgets tend to have restricted uses and software payments fit into very few categories. Effectively, I'd have to ask for a brand new gr…
I am starting to see folks saying - ok, so LLMs can do this, what value have you added ? modulo llm is becoming the norm.
Re: A recent experience with ChatGPT 5.5 Pro
#513It's a very long post with a mix of technical (math) and philosophical sections. Here are the most striking points to reflect upon IMHO. > It seems to me that training beginning PhD students to do research [...] has just got harder, since one obvious way to help somebody get started is to give them a problem that looks as though it might be a relatively gentle one. If LLMs are at the point where they can solve “gentl…
> by solving hard problems you get an insight into the problem-solving process itself, at least in your area of expertise, in a way that you simply don’t if all you do is read other people’s solutions. One consequence of this is that people who have themselves solved difficult problems are likely to be significantly better at using solving problems with the help of AI, just as very good coders are better at vibe codi…
I don't know or understand the binary executable.
I don't own the binary executable, I don't understand it, I can't explain it, it's not my work.
I'm a passthrough; I'm invisible.
I have literally no value whatsoever.
Re: A recent experience with ChatGPT 5.5 Pro
#514>> All I did was say things like, "Yes, it would be great if you could explore that idea and see whether you can get it to work," or "Could you rewrite that argument as a LaTeX file in the style of a standard mathematical preprint?"
Yeah, so all he did was take the horse to the water and make the horse drink. The collaboration with the other two mathematicians wasn't a trivial part of the problem solving either: every time Timothy Gowers figured ChatGPT had goe somewhere with its problem-solving, he stopped, asked it to render the answer in LaTex, and sent the answer off to be verified by the other two.
The reason for that is not to be underestimated: ChatGPT can produce answers to questions you ask it for as long as you ask it to do so but it has no capability to determine whether an answer is correct or not. That's why it needs a human with domain expertise to evaluate those answers. And of course to discard wrong answers in the process, because of course the process that's described here glosses over many false starts and back-and-forths and "you're absolutely rights, here's a new version of that"'s etc. that are common experience when using LLMs for problem-solving tasks.
The existential questions that the article poses about mathematics then are easily answered by taking all of the above into account. If LLMs are a useful tool for mathematicians, then nothing changes. Mathematicians of all levels can still do their job and perhaps do it faster or better with the new tool.
If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening.
Re: A recent experience with ChatGPT 5.5 Pro
#515This jives with what I've experienced in the brief time I had access to 5.5 Pro. It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. The downside (not noted in the…
> It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. I swear that people have said the same thing with effectively every new model that came out in the last six mo…
Re: A recent experience with ChatGPT 5.5 Pro
#516This jives with what I've experienced in the brief time I had access to 5.5 Pro. It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. The downside (not noted in the…
> can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, I don’t know about the rest of y’all but I find “rigidly guiding” LLM’s incredibly tedious and frustrating in the same way seeing an error code throw for the 40th time while troubleshooting something on my computer for two hours is frustrating. It also feels somewhat like…
Re: A recent experience with ChatGPT 5.5 Pro
#517Earlier quoted context omitted.
> The duo had jump-started the AI-for-Erdős craze late last year by prompting a free version of ChatGPT with open problems chosen at random from the Erdős problems website. (An AI researcher subsequently gifted them each a ChatGPT Pro subscription to encourage their “vibe mathing.”) Wonder who the AI researcher worked for? Is a "craze" something which a for-profit company would want to encourage? Maybe they'd think t…
OK, so let me get this straight. First, you ask for evidence of someone who isn't a VIP doing a similarly difficult problem using an LLM, to show that it isn't just VIPs being given special models. And then, when I provide that example, you say it doesn't count because the whole craze was started by researchers working for AI anyway. Furthermore, you start out stating that the access being given to these VIPs is to i…
How does any of this offend a reasonable person? It doesn't, because here is the project wiki addressing BOTH of these relevant points. https://github.com/teorth/erdosproblems/wiki/How-have-AI-com... https://github.com/teorth/erdosproblems/wiki/AI-contribution...
What I raised is a real question: the erdos folks are not affiliated with AI companies.. but is the AI company affiliated with them? Actually knowing the user/org accounts involved is optional because they could just divert resources to anyone prompting near the problem. Anecdotally.. I've noticed what appears to be token-discounts based on topic, for example more generosity for AI-related research than random stuff, but it's hard to know for sure. Wouldn't you promote interactions you could profitably train on?
So again, allocating resources to Erdos one way or another is just a clearly smart business decision for something people are talking about and which has become an unofficial competition among vendors, not a big scandalous accusation. Something like a reasoning-trace is the only way to settle it. This isn't conspiracy or nitpicking because this is the topic itself: the AI usage is more a matter of public interest than the actual problem solution. What's the argument against more transparency?
Re: A recent experience with ChatGPT 5.5 Pro
#518Earlier quoted context omitted.
What? DeepSeekV3 just came out and is incredible for the price. Mythos is also half-released.
Until you or I can actually use Mythos in Claude without an nda or other strings attached, Mythos is not released and is just an effective marketing tool for Anthropic.
Re: A recent experience with ChatGPT 5.5 Pro
#519Earlier quoted context omitted.
> "After 16 minutes and 41 seconds, it came back" ... "further 47 minutes and 39 seconds" ... "After 13 minutes and 33 seconds" ... "After 9 minutes and 12 seconds" ... "After 31 minutes and 40 seconds" ... plus other computations Anyone spotting the issue here? What did that really cost? Whatever the Joules... (convert to $ using your preferred benchmark price) it is a fraction to what it might take a human Ph. D. w…
Not necessarily. Humans brains use a tiny amount of power. Most of the human cost would be due to the very high cost of housing in many locations.
The real power required to support a human life in a developed country is a lot. Wattage for the human brain is definitely miniscule in comparison.
Re: A recent experience with ChatGPT 5.5 Pro
#520Earlier quoted context omitted.
> by solving hard problems you get an insight into the problem-solving process itself, at least in your area of expertise, in a way that you simply don’t if all you do is read other people’s solutions. One consequence of this is that people who have themselves solved difficult problems are likely to be significantly better at using solving problems with the help of AI, just as very good coders are better at vibe codi…
Are you a cutting edge research scientist or something? Everyone I know works in the same domain every day. The problems are the same. People aren't solving brand new problems to humanity every day. We make budgets and look at ticket counts. Roll out patches. Replace hardware. Upgrade software packages. Make a new dashboard to track a project. I guess if every day is a completely novel thing for you, ok. I feel like…
My point is, if you delegate your job to AI, and it works, then 1/ you don't know the result of the work in more detail than any other person, and 2/ the people you're reporting to can probably write a prompt as good as yours, if not better.
Which means: you've made yourself dispensable. Nothing very good for dinner; no nice place to live. But lots of time to practice guitar I guess.