Live data from Hacker News

GPT-5.2

openai.com

481–490 of 1001 posts

Re: GPT-5.2

#481

Earlier quoted context omitted.

I think Grok's voice chat is almost there - only things missing for me: * it's slower to start-up by a couple of seconds * it's harder to switch between voice and text and back again in the same chat (though ChatGPT isn't perfect at this either) And of course Grok's unhinged persona is... something else.

Pretty good until it goes crazy glazing Elon or declaring itself mecha hitler.

Neither of these have happened in my use. Those were both the product of some pretty aggressive prompting, and were remedied months ago.

Re: GPT-5.2

#483
post #63
post #10

Earlier quoted context omitted.

I'm quite sad about the S-curve hitting us hard in the transformers. For a short period, we had the excitement of "ooh if GPT-3.5 is so good, GPT-4 is going to be amazing! ooh GPT-4 has sparks of AGI!" But now we're back to version inflation for inconsequential gains.

I don't feel the S-curve at all yet. Still an exponential for me

With a very long doubling time?

Re: GPT-5.2

#484

Earlier quoted context omitted.

I will say that it is wild, if not somewhat problematic that two users have such disparate views of seemingly the same product. I say that, but then I remember my own experience just from few days ago. I don't pay for gemini, but I have paid chatgpt sub. I tested both for the same product with seemingly same prompt and subbed chatgpt subjectively beat gemini in terms of scope, options and links with current decent de…

I can use GPT one day and the next get a different experience with the same problem space. Same with Gemini.

This is by design, given a non-determenitisic application?

Re: GPT-5.2

#485
post #434

Earlier quoted context omitted.

This is also the default in Gemini pretty sure, at least I remember turning it off. Make's no sense to me why this is the default.

> Makes no sense to me why this is the default. You’re probably pretty far from the average user, who thinks “AI is so dumb” because it doesn’t remember what you told it yesterday.

I was thinking more people would be annoyed by it bringing up unrelated conversations, thinking more I'd say you're probably right that more people are expecting it to remember everything they say.

Re: GPT-5.2

#486

Weirdly, the blog announcement completely omits the actual new context window size which is 400,000: https://platform.openai.com/docs/models/gpt-5.2 Can I just say !!!!!!!! Hell yeah! Blog post indicates it's also much better at using the full context. Congrats OpenAI team. Huge day for you folks!! Started on Claude Code and like many of you, had that omg CC moment we all had. Then got greedy. Switched over to Codex…

I haven't done a ton of testing due to cost, but so far I've actually gotten worse results with xhigh than high with gpt-5.1-codex-max. Made me wonder if it was somehow a PEBKAC error. Have you done much comparison between high and xhigh?

For a few weeks the Codex model has been cursed. Recommend sticking with 5.1 high , 5.2 feels good too but early days

Re: GPT-5.2

#487

Earlier quoted context omitted.

I will say that it is wild, if not somewhat problematic that two users have such disparate views of seemingly the same product. I say that, but then I remember my own experience just from few days ago. I don't pay for gemini, but I have paid chatgpt sub. I tested both for the same product with seemingly same prompt and subbed chatgpt subjectively beat gemini in terms of scope, options and links with current decent de…

It’s like having 3 coins and users preferring one or the other when tossing it because one coin gives consistently more heads (or tails) than the other coin. What is better is to build a good set of rules and stick to one and then refine those rules over time as you get more experience using the tool or if the tool evolves and digress from the results you expect.

But, unless you are on a local model you control, you literally can't. Otherwise, good rules will work only as long as the next update allows. I will admit that makes me consider some other options, but those probably shouldn't be 'set and iterate' each time something changes.

Re: GPT-5.2

#488

Earlier quoted context omitted.

Last year o3 high did 88% on ARC-AGI 1 at more than $4,000/task. This model at its X high configuration scores 90.5% at just $11,64 per task. General intelligence has ridiculously gotten less expensive. I don't know if it's because of compute and energy abundance,or attention mechanisms improving in efficiency or both but we have to acknowledge the bigger picture and relative prices.

Sure, but the reason I'm confused by the pricing is that the pricing doesn't exist in a vacuum. Pro barely performs better than Thinking in OpenAI's published numbers, but comes at ~10x the price with an explicit disclaimer that it's slow on the order of minutes. If the published performance numbers are accurate, it seems like it'd be incredibly difficult to justify the premium. At least on the surface level, it look…

It could be using the same early trick of Grok (at least in the earlier versions) that they boot 10 agents who work on the problem in parallel and then get a consensus on the answer. This would explain the price and the latency.

Essentially a newbie trick that works really well but not efficient, but still looking like it's amazing breakthrough.

(if someone knows the actual implementation I'm curious)

Re: GPT-5.2

#490
post #311
post #31

For me the last remaining killer feature of ChatGPT is the quality of the voice chat. Do any of the competitors have something like that?

On the contrary, I thought Gemini 3 Live mode is much much better than ChatGPT. The voices have none of the annoying artificial uptalking intonations that ChatGPT has, and the simplex/duplex interruptibility of Gemini Live seems more responsive. It knows when to break and pause during conversations.

Apart from sounding a bit stiff and informal, I was also surprised at how good Gemini Live mode is in regional Indian languages.
Post reply on HN