Earlier quoted context omitted.
Apparently they have not had a successful pre training run in 1.5 years
I want to read a short scify story set in 2150 about how, mysteriously, no one has been able to train a better LLM for 125 years. The binary weights are studied with unbelievably advanced quantum computers but no one can really train a new AI from scratch. This starts cults, wars and legends and ultimately (by the third book) leads to the main protagonist learning to code by hand, something that no human left alive s…
GPT-5.2
91–100 of 1001 posts
Re: GPT-5.2
#92From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!
That ARC AGI score is a little suspicious. That's a really tough for AI benchmark. Curious if there were improvements to the test harness because that's a wild jump in general problem solving ability for an incremental update.
Still waiting of Full Self Driving myself.
Re: GPT-5.2
#93An almost 50% price increase. Benchmarks look nice, but 50% more nice...?
Re: GPT-5.2
#94Earlier quoted context omitted.
We're also in benchmark saturation territory. I heard it speculated that Anthropic emphasizes benchmarks less in their publications because internally they don't care about them nearly as much as making a model that works well on the day-to-day
How do you measure whether it works better day to day without benchmarks?
Re: GPT-5.2
#95"Investors are putting pressure, change the version number now!!!"
I'm quite sad about the S-curve hitting us hard in the transformers. For a short period, we had the excitement of "ooh if GPT-3.5 is so good, GPT-4 is going to be amazing! ooh GPT-4 has sparks of AGI!" But now we're back to version inflation for inconsequential gains.
Re: GPT-5.2
#96Everything is still based on 4 4o still right? is a new model training just too expensive? They can consult deepseek team maybe for cost constrained new models.
Where did you get that from? Cutoff date says august 2025. Looks like a newly pretrained model
Re: GPT-5.2
#97From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!
Re: GPT-5.2
#98Re: GPT-5.2
#99This seems like another "better vibes" release. With the number of benchmarks exploding, random luck means you can almost always find a couple showing what you want to show. I didn't see much concrete evidence this was noticeably better than 5.1 (or even 5.0). Being a point release though I guess that's fair. I suspect there is also some decent optimizations on the backend that make it cheaper and faster for OpenAI t…
> I didn't see much concrete evidence this was noticeably better than 5.1 Did you test it?
Re: GPT-5.2
#100Earlier quoted context omitted.
Apparently they have not had a successful pre training run in 1.5 years
I want to read a short scify story set in 2150 about how, mysteriously, no one has been able to train a better LLM for 125 years. The binary weights are studied with unbelievably advanced quantum computers but no one can really train a new AI from scratch. This starts cults, wars and legends and ultimately (by the third book) leads to the main protagonist learning to code by hand, something that no human left alive s…
Might sell better with the protagonist learning iron age leatherworking, with hides tanned from cows that were grown within earshot, as part of a process of finding the real root of the reason for why any of us ever came to be in the first place. This realization process culminates in the formation of a global, unified steampunk BDSM movement and a wealth of new diseases, and then: Zombies.
(That's the end. Zombies are always the end.)