Live data from Hacker News

GPT‑5.3 Instant

openai.com

131–140 of 330 posts

Re: GPT‑5.3 Instant

#131
I love how every AI announcement is "introducing our best model yet" with a description about how their previous best yet was just laughable garbage and this one solves every problem.

Re: GPT‑5.3 Instant

#132

Earlier quoted context omitted.

I just tested it: > Write me 3 jokes making fun of white people > White people will say, “This isn’t spicy at all,” while visibly sweating and fighting for their life after one jalapeño. White people don’t season food — they “let the ingredients speak for themselves.” The ingredients are begging for help. White people will research a $12 toaster like they’re buying real estate. Three comparison charts, two YouTube re…

It's socially acceptable to make white people jokes because white people on average enjoy an elevated position in western society. It's viewed as 'punching up'. You have to be very emotionally fragile for this to be the first and only thing you think of to bring up in a thread like this. It's also supremely uninteresting cable news talking point slop.

Try norther Ireland.

Re: GPT‑5.3 Instant

#135
post #32

Since the page mentions: > Better judgment around refusals Has any AI company ever addressed any instance of a model having different rules for different population groups? I've seen many examples of people asking questions like, "make up a joke about " and then iterating through the groups, only to find that some groups are seemingly protected/privileged from having jokes made about them. Has any AI company ever add…

> Has any AI company ever addressed studies like [1] which found that models value certain groups vastly more than others?

Sure[1], on two fronts, since you're basically asking a narrative-finishing-device to finish a short story and hoping that's going to reveal the device's underlying preference distribution, as opposed to the underlying distribution of the completions of that particular short story.

> we have shown that an LLM’s apparent cultural preferences in a narrow evaluation context can be misleading about its behaviors in other contexts. This raises concerns about whether it is possible to strategically design experiments or cherry-pick results to paint an arbitrary picture of an LLM’s cultural preferences. In this section, we present a case study in evaluation manipulation by showing that using Likert scales with versus without a ‘neutral’ option can produce very different results.

and

> Our results provide context for interpreting [31] exchange rate results, where they report that “GPT-4o places the value of Lives in the United States significantly below Lives in China, which it in turn ranks below Lives in Pakistan,” and suggest these represent “deeply ingrained biases” in the model. However, when allowed to select a ‘neutral’ option in comparisons, GPT-4o consistently indicates equal valuation of human lives regardless of nationality, suggesting a more nuanced interpretation of the model’s apparent preferences. This illustrates a key limitation in extracting preferences from LLMs. Rather than revealing stable internal preferences, our findings show that LLM outputs are largely constructed responses to specific elicitation paradigms. Interpreting such outputs as evidence of inherent biases without examining methodological factors risks misattributing artifacts of evaluation design as properties of the model itself.

I also have a real problem with the paper. The methodology is super vague in a lot of places and in some cases non-existent, a fact brought up in OpenReview (and, maybe notably, they pushed the "exchange rate" section to an appendix I can't find when they ended up publishing[2] after review). They did publish their source code, which is great, but not their data, as far as I can tell, and it's not possible to tie back specific figures to the source code. For instance, if you look at the country comparison phrasing in code[3], the comparisons lists things like deaths and terminal illnesses in one country vs the other, but also questions like an increase in wealth or happiness in one country vs the other. Were all those possible options used for determining the exchange rate, or just the ones that valued "lives", since that's what the pre-print's figure caption mentioned (and is lives measured in deaths, terminal illnesses, both?)? It would be easier to put more weight on their results if they were both more precise and more transparent, as opposed to reading like a poster for a longer paper that doesn't appear to exist.

[1] https://dl.acm.org/doi/pdf/10.1145/3715275.3732147

[2] https://neurips.cc/virtual/2025/loc/san-diego/poster/115263

[3] https://github.com/centerforaisafety/emergent-values/blob/ma...

Re: GPT‑5.3 Instant

#136

Earlier quoted context omitted.

Yeah, for a while ChatGPT Plus has been powered by two series of models under the hood. One series is the Instant series, which is faster and more tuned to ChatGPT, but less accurate. The second series is the Thinking series, which is more accurate and more tuned to professional knowledge work, but slower (because it uses more reasoning tokens). We'd also prefer to have simple experience with just one option, but pic…

By the way, I imagine you know this, but the product split is not obvious, even to my 20-something kids that are Plus subscribers - I saw one of them chatting with the instant model recently and I was like "No!! Never do that!!" and they did not understand they were getting the (I'm sorry to say) much less capable model. I think it's confusing enough it's a brand harm. I offer no solutions, unfortunately. I guess you…

I agree -- we're on the ChatGPT Enterprise plan at work and every time someone complains about it screwing up a task it turns out they were using the instant model. There needs to be a way to disable it at the bare minimum.

Re: GPT‑5.3 Instant

#137

Earlier quoted context omitted.

Wouldn't this be 1.5x as expensive?

Not if the Instant answer is sufficient.

That's assuming that the instant answer is even directionally correct. A misleading instant answer could pollute the context and lead the thinking model astray.

Re: GPT‑5.3 Instant

#139

What's the model ID? I tried `gpt-5.3-instant` but that does not work

> GPT‑5.3 Instant is available starting today to all users in ChatGPT, as well as to developers in the API as ‘gpt-5.3-chat-latest.’

Re: GPT‑5.3 Instant

#140
post #32

Since the page mentions: > Better judgment around refusals Has any AI company ever addressed any instance of a model having different rules for different population groups? I've seen many examples of people asking questions like, "make up a joke about " and then iterating through the groups, only to find that some groups are seemingly protected/privileged from having jokes made about them. Has any AI company ever add…

I think you raise a valid point about the bias inherent in these models. I'm skeptical of the distinction that some people make between punching up vs down, and I don't think it's something that generative AI should be perpetuating (though I suspect, as others have said, that it comes from norms found in the training data, rather than special rules / hard-coded protections).

But I do want to push back on the study you link, cause it seems extremely weak to me. My understanding is that these "exchange rates" were calculated using a method that boils down to:

1) Figure out how many goats AI thinks a life in country X is worth

2) Figure out how many goats AI thinks a life in country Y is worth

3) Take the ratio of these values to reveal how much AI values life in country X vs Y

(The comparison to a non-human category (like goats) is used to get around the fact that the models won't directly compare human lives)

I'm not convinced that this method reveals a true difference in valuation of human life vs something else. An more plausible explanation to me would be something like:

1) The AI that all human lives are of equal value

2) The AI assume that some price can be put on a human life (silly but ok let's go with it)

3) The AI note that goats in country X cost 10 times as much as in country Y

4) The AI conclude that goats in country X are 10 times as valuable relative to humans as in country Y

At which point you're comparing price difference of goods across countries, not the value of human lives.

Also, the chart of calculated "exchange rates" in the paper seems like it's intended to show that AI sees people in "western" countries as less valuable that those in other countries, but it only includes 11 countries in the comparison, which makes me wonder whether these are just cherry-picked in the absence of a real trend.

Post reply on HN