Live data from Hacker News

GPT-5.2

openai.com

91–100 of 1001 posts

Re: GPT-5.2

#91
post #9

Earlier quoted context omitted.

Apparently they have not had a successful pre training run in 1.5 years

I want to read a short scify story set in 2150 about how, mysteriously, no one has been able to train a better LLM for 125 years. The binary weights are studied with unbelievably advanced quantum computers but no one can really train a new AI from scratch. This starts cults, wars and legends and ultimately (by the third book) leads to the main protagonist learning to code by hand, something that no human left alive s…

There's a scifi short story about a janitor who knows how to do basic arithmetic and becomes the most important person in the world when some disaster happens. Of course after things get set up again due to his expertise, he becomes low status again.

Re: GPT-5.2

#92
post #47

From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!

That ARC AGI score is a little suspicious. That's a really tough for AI benchmark. Curious if there were improvements to the test harness because that's a wild jump in general problem solving ability for an incremental update.

I don’t think their words mean just about anything, only the behavior of the models.

Still waiting of Full Self Driving myself.

Re: GPT-5.2

#93

An almost 50% price increase. Benchmarks look nice, but 50% more nice...?

#1 models are usually priced at 2x more than the competition, and they often decrease the price right when they lose the crown.

Re: GPT-5.2

#94

Earlier quoted context omitted.

We're also in benchmark saturation territory. I heard it speculated that Anthropic emphasizes benchmarks less in their publications because internally they don't care about them nearly as much as making a model that works well on the day-to-day

How do you measure whether it works better day to day without benchmarks?

Internal evals, Big AI certainly has good, proprietary training and eval data, it's one reason why their models are better

Re: GPT-5.2

#95
post #10
post #5

"Investors are putting pressure, change the version number now!!!"

I'm quite sad about the S-curve hitting us hard in the transformers. For a short period, we had the excitement of "ooh if GPT-3.5 is so good, GPT-4 is going to be amazing! ooh GPT-4 has sparks of AGI!" But now we're back to version inflation for inconsequential gains.

Because it will take thousands of underpaid researchers random searching through solution space to get to the next improvement, not 2-3 companies pressed to monetize and enshittify their product before money runs out. That and winning more hardware lotteries.

Re: GPT-5.2

#96
post #3

Everything is still based on 4 4o still right? is a new model training just too expensive? They can consult deepseek team maybe for cost constrained new models.

Where did you get that from? Cutoff date says august 2025. Looks like a newly pretrained model

If the pretraining rumors are true, they're probably using continued pretraining on the older weights. Right?

Re: GPT-5.2

#97

From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!

Open AI has already been busted for getting benchmark information and training the models on that. At this point if you believe Sam Altman, I have a bridge to sell you.

Re: GPT-5.2

#99
post #24

This seems like another "better vibes" release. With the number of benchmarks exploding, random luck means you can almost always find a couple showing what you want to show. I didn't see much concrete evidence this was noticeably better than 5.1 (or even 5.0). Being a point release though I guess that's fair. I suspect there is also some decent optimizations on the backend that make it cheaper and faster for OpenAI t…

> I didn't see much concrete evidence this was noticeably better than 5.1 Did you test it?

No, I would like to but I don't see it in my paid ChatGPT plan or in the API yet. I based my comment solely off of what I read in the linked announcement.

Re: GPT-5.2

#100
post #9

Earlier quoted context omitted.

Apparently they have not had a successful pre training run in 1.5 years

I want to read a short scify story set in 2150 about how, mysteriously, no one has been able to train a better LLM for 125 years. The binary weights are studied with unbelievably advanced quantum computers but no one can really train a new AI from scratch. This starts cults, wars and legends and ultimately (by the third book) leads to the main protagonist learning to code by hand, something that no human left alive s…

Sounds good.

Might sell better with the protagonist learning iron age leatherworking, with hides tanned from cows that were grown within earshot, as part of a process of finding the real root of the reason for why any of us ever came to be in the first place. This realization process culminates in the formation of a global, unified steampunk BDSM movement and a wealth of new diseases, and then: Zombies.

(That's the end. Zombies are always the end.)

Post reply on HN