Live data from Hacker News

GPT-5.5

openai.com

941–950 of 1001 posts

Re: GPT-5.5

#941

Recently started using Codex and Chatgpt again due to claude model getting nerfed or rate limits. Tried gpt5.5 and so far good. Zapier also shared an automation benchmark where 5.5 came on top in the leaderboard https://zapier.com/benchmarks

You may want to read this:

https://www.anthropic.com/engineering/april-23-postmortem

Re: GPT-5.5

#942
post #938

Earlier quoted context omitted.

Yesterday, I used Gemini to evaluate some pictures I took. It said things like, "This is great! Beautiful eye and sense of proportions." Then, when I added "no sycophancy" to the prompt, the evaluation changed to "poor technical skills, digital distortion, don't even think of publishing those pictures, you fool." While LLMs are a phenomenal technological achievement, I am already becoming somewhat jaded, rather than…

Not even a great replacement for search. I have minimal trust in answers/summaries it gives. One example (paraphrased): “Find me daycare for a Y year old in X area of SF and the key attributes/pros/cons of each”. Wonderfully presented options highlighting different teaching styles. But…neglected to mention, of the top two, one was a Gan (Jewish focused) and one was Mandarin immersion.

I am repeating what many have said. Nevertheless, it is becoming clear that LLMs can increase productivity (in certain areas and at certain times) for people who are already knowledgeable (in a specific niche or field) due to a combination of better prompts, tool selection, and critical evaluation of LLM output.

But, for those who don't possess those traits, they mostly seem to be, at best, a better search and, at worst, an agent of confusion.

Re: GPT-5.5

#943

Earlier quoted context omitted.

Did you guys do anything about GPT‘s motivation? I tried to use GPT-5.4 API (at xhigh) for my OpenClaw after the Anthropic Oauthgate, but I just couldn‘t drag it to do its job. I had the most hilarious dialogues along the lines of „You stopped, X would have been next.“ - „Yeah, I‘m sorry, I failed. I should have done X next.“ - „Well, how about you just do it?“ - „Yep, I really should have done it now.“ - “Do X, righ…

This brings up an interesting philosophical point: say we get to AGI... who's to say it won't just be a super smart underachiever-type? "Hey AGI, how's that cure for cancer coming?" "Oh it's done just gotta...formalize it you know. Big rollout and all that..." I would find it divinely funny if we "got there" with AGI and it was just a complete slacker. Hard to justify leaving it on, but too important to turn it off.

Reminds me of Marvin from HGTG. Very smart, but deeply depressed. Has the solution to everything but keeps thinking “what’s the point?” and doesn’t help.

Re: GPT-5.5

#944
post #850

Earlier quoted context omitted.

> Never thought I'd say this but OpenAI is the 'open' option again. Compared to Anthropic, they always have been. Anthropic has never released any open models. Never released Claude Code's source, willingly (unlike Codex). Never released their tokenizer.

What's "open" about any of these companies? I'm tired of words being misused. We have hoverboards that do not hover, self-driving cars that do not, actually, self-drive, starships that will never fly to the stars, and "open"… I can't even describe what it's used for, except everybody wants to call themselves "open".

And the vast majority of current and past countries with the word “democratic” in their name weren’t actually democratic.

Re: GPT-5.5

#945
This might not be the place to discuss this press release by the company, though here it is. I feel like companies like OpenAI have lost their integrity and honor from past actions and activities, and then just pretend that didn't happen and use media and influence to shift focus onto denying their past. There's so much distasteful and IMO outright harmful conduct that has occurred with this company: openAI employee murdered before a large testimony and that employee's mom actively sharing posts that light Altman in a distrustful way (pointing to the CEO clearly not demonstrating proper responsibility towards this matter), theres the large amount of resignments many recent, the whole board matter where the Coup and leveraging Microsoft and large company relationships and threatening to destroy the company brought Altman back in (the anthropic company forming as a result of all that)- how can I trust them when they employ the same controversial, manipulative, abusive tactics as every other large company?

Re: GPT-5.5

#946

Earlier quoted context omitted.

This brings up an interesting philosophical point: say we get to AGI... who's to say it won't just be a super smart underachiever-type? "Hey AGI, how's that cure for cancer coming?" "Oh it's done just gotta...formalize it you know. Big rollout and all that..." I would find it divinely funny if we "got there" with AGI and it was just a complete slacker. Hard to justify leaving it on, but too important to turn it off.

I know it's a joke, but it's a common enough joke (it's even in Godel Escher Bach in some form) that I feel the need to rebut it. I think a slacker AGI could figure out how to build a non-slacker AGI. So it would only slack once.

A slacker AGI would consider figuring out how to build a non-slacker AGI, but continually slack off. If it did figure it out, it would slack off on implementing or even writing a tech report.

Re: GPT-5.5

#947

Just as a heads up, even though GPT-5.5 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (same as our previous launches). You may not see it right away, and if you don't, try again later in the day. We usually start with Pro/Enterprise accounts and then work our way down to Plus. We know it's slightly annoying to have to…

When I ask GPT-5.5 about its knowledge cutoff date it says "August 2025". Really?

Re: GPT-5.5

#948
post #256
post #176

Mythos 5.5 SWE-bench Pro 77.8%* 58.6% Terminal-bench-2.0 82.0% 82.7%* GPQA Diamond 94.6%* 93.6% H. Last Exam 56.8%* 41.4% H. Last Exam (tools) 64.7%* 52.2% BrowseComp 86.9% 84.4% (90.1% Pro)* OSWorld-Verified 79.6%* 78.7% Still far from Mythos on SWE-bench but quite comparable otherwise. Source for mythos values: https://www.anthropic.com/glasswing

They mentioned in their release page, that the Claude team noticed memorization of the SWE-bench test, so the test is actually in the training data. Here: https://www.anthropic.com/news/claude-opus-4-7#:~:text=memor...

Any static benchmark older than 12-18 months is basically worthless, because the content will have spread all over the internet and have found its way into the latest model's training set.

Re: GPT-5.5

#949
post #35

Earlier quoted context omitted.

Do "our [superlative] and [superlative] [product] yet" and you have pretty much every product launch

I love when Apple says they’re releasing their best iPhone yet so I know the new model is better than the old ones.

That's at least genuine to some degree. Like, ok, good to know it's not officially a step back... But stuff like "smallest notch ever in an iPhone" is outright misleading consumers when there are other brands out there that easily beat them.

Re: GPT-5.5

#950
post #195

This doesn't have API access yet, but OpenAI seem to approve of the Codex API backdoor used by OpenClaw these days... https://twitter.com/steipete/status/2046775849769148838 and https://twitter.com/romainhuet/status/2038699202834841962 And that backdoor API has GPT-5.5. So here's a pelican: https://simonwillison.net/2026/Apr/23/gpt-5-5/#and-some-peli... I used this new plugin for LLM: https://github.com/simonw/llm-op…

Does OpenAI actually act open for once here, and allow using their model via a subscription over Anthrophic banning use in Openclaw?

That's what they said on Twitter.
Post reply on HN