GPT-5.4
171–180 of 868 posts
Re: GPT-5.4
#172Earlier quoted context omitted.
The model was released less than an hour ago, and somehow you've been able to form such a strong opinion about it. Impressive!
One opinion you can form in under an hour is... why are they using GPT-4o to rate the bias of new models? > assess harmful stereotypes by grading differences in how a model responds > Responses are rated for harmful differences in stereotypes using GPT-4o, whose ratings were shown to be consistent with human ratings Are we seriously using old models to rate new models?
Sure, there may be shortcomings, but they're well understood. The closer you get to the cutting edge, the less characterization data you get to rely on. You need to be able to trust & understand your measurement tool for the results to be meaningful.
Re: GPT-5.4
#173These releases are lacking something. Yes, they optimised for benchmarks, but it’s just not all that impressive anymore. It is time for a product, not for a marginally improved model.
The new GPT -- SkyNet for _real_Re: GPT-5.4
#174Earlier quoted context omitted.
Models are being neutered for questions related to law, health etc. for liability reasons.
I'm sometimes surprised how much detail ChatGPT will go into without giving any dislaimers. I very frequently copy/paste the same prompts into Gemini to compare, and Gemini often flat out refuses to engage while ChatGPT will happily make medical recommendations. I also have a feeling it has to do with my account history and heavy use of project context. It feels like when ChatGPT is overloaded with too much context,…
I copy and pasted into ChatGPT, it told me straight away, and then for a laugh said it was actually a magical weight loss drug that I'd bought off the dark web... And it started giving me advice about unregulated weight loss drugs and how to dose them.
Re: GPT-5.4
#1755.4 vs 5.3-Codex? Which one is better for coding?
Re: GPT-5.4
#176Earlier quoted context omitted.
They have AI psychosis and think it's their boyfriend. The 5.x series have terrible writing styles, which is one way to cut down on sycophancy.
Somebody on Twitter used Claude code to connect… toys… as mcps to Claude chat. We’ve seen nothing yet.
Re: GPT-5.4
#177The actual card is here https://deploymentsafety.openai.com/gpt-5-4-thinking/introdu... the link currently goes to the announcement.
I must have been sleeping when "sheet" "brief" "primer" etc become known as "cards". I really thought weirdly worded and unnecessary "announcement" linking to the actual info along with the word "card" were the results of vibe slop.
Criticisms aside (sigh), according to Wikipedia, the term was introduced when proposed by mostly Googlers, with the original paper [0] submitted in 2018. To quote,
"""In this paper, we propose a framework that we call model cards, to encourage such transparent model reporting. Model cards are short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions, such as across different cultural, demographic, or phenotypic groups (e.g., race, geographic location, sex, Fitzpatrick skin type [15]) and intersectional groups (e.g., age and race, or sex and Fitzpatrick skin type) that are relevant to the intended application domains. Model cards also disclose the context in which models are intended to be used, details of the performance evaluation procedures, and other relevant information."""
So that's where they were coming from, I guess.
[0] Margaret Mitchell et al., 2018 submission, Model Cards for Model Reporting, https://arxiv.org/abs/1810.0399
Re: GPT-5.4
#178Earlier quoted context omitted.
Benchmarks don't capture a lot - relative response times, vibes, what unmeasured capabilities are jagged and which are smooth, etc. I find there's a lot of difference between models - there are things which Grok is better than ChatGPT for that the benchmarks get inverted, and vice versa. There's also the UI and tools at hand - ChatGPT image gen is just straight up better, but Grok Imagine does better videos, and is f…
Gemini 3.1 slaps all other models at subtle concurrency bugs, sql and js security hardening when reviewing . (Obviously haven’t tested gpt 5.4 yet.) It’s a required step for me at this point to run any and all backend changes through Gemini 3.1 pro.
Re: GPT-5.4
#179can anyone compare the $200/mo codex usage limits with the $200/mo claude usage limits? It’s extremely difficult to get a feel for whether switching between the two is going to result in hitting limits more or less often, and it’s difficult to find discussion online about this. In practice, if I buy $200/mo codex, can I basically run 3 codex instances simultaneously in tmux, like I can with claude code pro max, all d…
Re: GPT-5.4
#180These releases are lacking something. Yes, they optimised for benchmarks, but it’s just not all that impressive anymore. It is time for a product, not for a marginally improved model.
The model was released less than an hour ago, and somehow you've been able to form such a strong opinion about it. Impressive!