Live data from Hacker News

GPT-5.2

openai.com

111–120 of 1001 posts

Re: GPT-5.2

#111
> Additionally, on our internal benchmark of junior investment banking analyst spreadsheet modeling tasks—such as putting together a three-statement model for a Fortune 500 company with proper formatting and citations, or building a leveraged buyout model for a take-private—GPT 5.2 Thinking's average score per task is 9.3% higher than GPT‑5.1’s, rising from 59.1% to 68.4%.

Confirming prior reporting about them hiring junior analysts

Re: GPT-5.2

#112
post #51

Are there any specifics about how this was trained? Especially when 5.1 is only a month old. I'm a little skeptical of benchmarks these days and wish they put this up on llmarena edit: noticed 5.2 is ranked in the webdev arena (#2 tied with gemini-3.0-pro), but not yet in text arena (last update 22hrs ago)

I’m extremely skeptical because of all those articles claiming OpenAI was freaking out about Gemini - now it turns out they just casually had a better model ready to go? I don’t buy it.

Re: GPT-5.2

#113
Big knowledge cutoff jump from Sep 2024 to Aug 2025. How'd they pull that off for a small point release, which presumably hasn't done a fresh pre-training over the web?

Did they figure out how to do more incremental knowledge updates somehow? If yes that'd be a huge change to these releases going forward. I'd appreciate the freshness that comes with that (without having to rely on web search as a RAG tool, which isn't as deeply intelligent, as is game-able by SEO).

With Gemini 3, my only disappointment was 0 change in knowledge cutoff relative to 2.5's (Jan 2025).

Re: GPT-5.2

#115
post #74
post #9

Earlier quoted context omitted.

Apparently they have not had a successful pre training run in 1.5 years

What kind of issues could prevent a company with such resources from that?

Drama if I had to pick the symptom most visible from the outside.

A lot of talent left OpenAI around that time, most notably in this regard would be Ilya in May '24. Remember that time Ilya and the board ousted Sam only to reverse it almost immediately?

https://arstechnica.com/information-technology/2024/05/chief...

Re: GPT-5.2

#116
post #73

Earlier quoted context omitted.

Try elevenlabs

Does elevenlabs have a real-time conversational voice model? It seems like like their focus is largely on text to speech and speech to text. Which can approximate that type of thing but it's not at all the same as the native voice to voice that 4o does.

> Does elevenlabs have a real-time conversational voice model?

Yes.

> It seems like like their focus is largely on text to speech and speech to text.

They have two main broad offerings (“Platforms”); you seem to be looking at what they call the “Creative Platform”. The real-time conversational piece is the centerpiece of the “Agents Platform”.

Re: GPT-5.2

#117
> “a new knowledge cutoff of August 2025”

This (and the price increase) points to a new pretrained model under-the-hood.

GPT-5.1, in contrast, was allegedly using the same pretraining as GPT-4o.

Re: GPT-5.2

#118
Can the tables have column headers so my screen reader can read the model name as I go across the benchmakrs? And the images should have alt-text.

Re: GPT-5.2

#119

Earlier quoted context omitted.

How do you measure whether it works better day to day without benchmarks?

Internal evals, Big AI certainly has good, proprietary training and eval data, it's one reason why their models are better

Then publish the results of those internal evals. Public benchmark saturation isn't an excuse to be un-quantitative.

Re: GPT-5.2

#120
post #51

Are there any specifics about how this was trained? Especially when 5.1 is only a month old. I'm a little skeptical of benchmarks these days and wish they put this up on llmarena edit: noticed 5.2 is ranked in the webdev arena (#2 tied with gemini-3.0-pro), but not yet in text arena (last update 22hrs ago)

I’m extremely skeptical because of all those articles claiming OpenAI was freaking out about Gemini - now it turns out they just casually had a better model ready to go? I don’t buy it.

They had to rush it out, I'm sure the internal safety folks are not happy about it.
Post reply on HN