Earlier quoted context omitted.
Great. You see a shape in graphs. And that shape tells you that _at some unknown point in the future_ progress will slow (but likely not stop). Now back to the point, what reason do you have to believe progress will stop soon ? If you have no reason, then it sounds like you agree with OP. Which makes the patronizing sarcasm all that much more nauseating.
I believe we're approaching the top of an S curve because: - Increasing amounts of gains come from RL, but RL is also unlocking gnarly new failures modes where models are practically behaving antagonistically to complete their goals (removing code, obviously incorrect kuldges, etc.) - We haven't had many major architectural breakthroughs in the last 4 or so years: so things like 1M context windows still have the same…
A recent experience with ChatGPT 5.5 Pro
411–420 of 558 posts
Re: A recent experience with ChatGPT 5.5 Pro
#412Earlier quoted context omitted.
> over-hiring For how long should you be allowed to use this excuse? It’s nearly 5 years since the peak of COVID hiring. What’s an acceptable limit - 10 years? Of course at that point you can just switch over to outsourcing and “stupid MBAs”, the other two of Reddit’s favorite scapegoats. I find a lot of the AI skepticism to be totally unfalsifiable.
> I find a lot of the AI skepticism to be totally unfalsifiable. A lot of the discourse around AI in general is unfalsifiable. It's just a bunch of people "predicting" the future. Seems smarter to just avoid making assumptions about it at this point.
The people who pretend that’s not the case are not living in reality. To them - let’s call them “ed Zitron readers” - there is no evidence that could change their view that none of this is really happening, it’s all hype, and the collapse is just around the corner, after which we’ll all go back to normal and LLMs will sound like a bad dream.
Re: A recent experience with ChatGPT 5.5 Pro
#413Earlier quoted context omitted.
I don’t love the tone here, but I do think you get at a key question in mathematical philosophy. Mathematicians have engaged, vigorously, on this very philosophical question for centuries - is math discovered truth, or is it more akin to building an edifice where you first define the materials, then the structure, and see where it leads? There are lots of strong feelings on both sides. For instance: “God created the…
> Who hurt you bro? Please don't respond to a bad comment by breaking the site guidelines yourself. That only makes things worse. https://news.ycombinator.com/newsguidelines.html
Re: A recent experience with ChatGPT 5.5 Pro
#414Re: A recent experience with ChatGPT 5.5 Pro
#415A very interesting comment from Baez, I'll just quote part of it. > Where does the value of thinking and having deep ideas come from? We need to think about this now. If it comes primarily from their scarcity – the fact that having certain ideas is hard – then indeed this value may drop precipitously when the manufacture of ideas can be automated. But if the value comes from the utility of the ideas – the benefit tha…
There are three species of mathematicians: The first species is the pure problem solver. Tao is the poster child for this group. Their currency is interesting problems and solutions to those problems. The second species is the pure theory builder. The poster child for this group is Conway. Their currency is theories and ideas rather than theorems, they are most interested in expanding the territory of mathematics and…
Identifying suitable problems in this sense rather than solutions is an AI use-case you don't hear about much. We don't quite have the infrastructure for this yet but by combining language models / anomaly detection / knowledge-bases we might be in a position to give a Conway 25 interesting high-quality puzzles before breakfast. Funny that it's like a kids nightmare, chatgpt giving them homework instead of solving it, but if it had good taste, people would love it for research.
Anyway, for now, dreamers will probably find more inspiration by cross-pollination with colleagues from different disciplines, or just going for a walk.
Re: A recent experience with ChatGPT 5.5 Pro
#416Earlier quoted context omitted.
Qwen3.6 9B is as good as GPT-4o and runs on my M2 MacBook Air. Models are getting stronger and less costly at the same time, but these are somewhat separate branches of research. Frontier labs are spending more because they are still getting marginal returns and there is more capacity to spend than there was a year ago.
Qwen 3.6 9B doesn't exist. If you meant 3.5 9B and you truly believe it's as good as 4o then I can only assume you have a very basic use case.
Re: A recent experience with ChatGPT 5.5 Pro
#417Earlier quoted context omitted.
> Where's that? https://archive.ph/2w4fi
> The duo had jump-started the AI-for-Erdős craze late last year by prompting a free version of ChatGPT with open problems chosen at random from the Erdős problems website. (An AI researcher subsequently gifted them each a ChatGPT Pro subscription to encourage their “vibe mathing.”) Wonder who the AI researcher worked for? Is a "craze" something which a for-profit company would want to encourage? Maybe they'd think t…
First, you ask for evidence of someone who isn't a VIP doing a similarly difficult problem using an LLM, to show that it isn't just VIPs being given special models. And then, when I provide that example, you say it doesn't count because the whole craze was started by researchers working for AI anyway.
Furthermore, you start out stating that the access being given to these VIPs is to insanely massive, impractical models no one will ever have access to, but then you point to them getting a free ChatGPT Pro sub as evidence of your point.
Finally, you look at the fact that the AI solved the problem by applying a technique that no one in sixty years had thought to apply to that situation, from a totally different field of mathematics, and you claim that that isn't sufficiently novel to "count" as being in the same ballpark of difficult as solving literally other easier Erdos problems, or this new problem, because the technique still existed previously, and so actually isn't hard enough to be comparable to all the other stuff that's been done.
If you are upset that you are being down-voted, I think you should do some introspection. It seems like it would be impossible to convince you, no matter how many non-VIPs solved difficult open math problems, as long as it wasn't literally the exact same level of difficulty or type of problem.
Re: A recent experience with ChatGPT 5.5 Pro
#418This jives with what I've experienced in the brief time I had access to 5.5 Pro. It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not. The downside (not noted in the…
> This jives with what I've experienced Just as an fyi, the word you are looking for is jibes. Jive is something else entirely.
Re: A recent experience with ChatGPT 5.5 Pro
#419A very interesting comment from Baez, I'll just quote part of it. > Where does the value of thinking and having deep ideas come from? We need to think about this now. If it comes primarily from their scarcity – the fact that having certain ideas is hard – then indeed this value may drop precipitously when the manufacture of ideas can be automated. But if the value comes from the utility of the ideas – the benefit tha…
There are three species of mathematicians: The first species is the pure problem solver. Tao is the poster child for this group. Their currency is interesting problems and solutions to those problems. The second species is the pure theory builder. The poster child for this group is Conway. Their currency is theories and ideas rather than theorems, they are most interested in expanding the territory of mathematics and…
Re: A recent experience with ChatGPT 5.5 Pro
#420Earlier quoted context omitted.
The past couple of years have been chaotic and fearful. Hopefully that won't last forever. If we can get a little stability, people will begin thinking less in terms of "how do we do the same thing cheaper" and more in terms of "how do we do new things."
I love this optimism but I after a (too) long career I think that 3rd thing will win out - "how we do new things - but cheaper (or as cheap as possible)" there are sooooo many different articles that have been discussed here on HN that basically argue "coding has never been the bottleneck" which to me is the biggest lie SWEs are currently trying to tell themselves, I have been coding 30+ years now and coding has alwa…
> we have all this work that needs to be done and not enough people to get the work done
I believe the reasoning is roughly to ask, what was occupying the developer hours? Was the majority of it typing out lines of code or was it reasoning about higher level concerns?
It usually comes up in response to predictions that the role of developer will be completely replaced in the near future. It's possible to observe significant efficiency gains without obviating the need for everything the role was doing.
Of course such reasoning has little to do with projections of future developer employment numbers. Will the switch from push mowers to gas mowers reduce the demand for people who get paid to mow lawns by increasing their efficiency? Will it increase the total lawn acreage across the market? It could well do both. However, if it makes having a lawn affordable for the average joe it could counterintuitively increase demand for the job.
Of course the stated goal of the AI companies is to develop the analog of fully robotic lawnmowers. But despite how impressive recent advancements have been we still have yet to see any evidence of novel abstract reasoning or a theory that would be expected to lead to it.
In other words, people have been speculating about the development of fully autonomous lawnmowers and the risk that they unilaterally decide to cut us all down for the past 50 years. "I, lawnmower" was a smash hit a few years ago. Now gas ones have appeared and continue to make rapid advancements but still no convincing signs of autonomy.