Live data from Hacker News

GPT-5.5

openai.com

961–970 of 1001 posts

Re: GPT-5.5

#961
post #305

Earlier quoted context omitted.

Can someone explain how we arrived at the pelican test? Was there some actual theory behind why it's difficult to produce? Or did someone just think it up, discover it was consistently difficult, and now we just all know it's a good test?

I set it up as a joke, to make fun of all of the other benchmarks. To my surprise it ended up being a surprisingly good measure of the quality of the model for other tasks (up to a certain point at least), though I've never seen a convincing argument as to why. I gave a talk about it last year: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ It should not be treated as a serious benchmark.

how can you say "it ended up being a surprisingly good measure of the quality of the model for other tasks" and also "It should not be treated as a serious benchmark" in the same comment?

if it is indeed a good measure of the quality of the model (hint: it's not) then, logically, it should be taken seriously.

this is, sadly, a great example of the kind of doublethink the "AI" hypesters (yes - whether you like it or not simon - that is what you are now) are all too capable of.

Re: GPT-5.5

#962
post #890

Earlier quoted context omitted.

I use pi.dev. I get openai team plan at work. Claude enterprise too. I have openrouter for myself. I use minimax 2.7. Kimi 2.6. And gpt 5.5 and opus 4.7. I can toggle between them in an open source interface that's how I stay able to not be trapped. Minimax is so cheap and for personal stuff it works fine. So I'm always toggling between the nre releases

what about just personal stuff in a syncing interface, what do you use for that?

What's a syncing interface?

Re: GPT-5.5

#963
post #762

I've found myself so deeply embedded in the Claude Max subscription that I'm worried about potentially makign a switch. How are people making sure they stay nimble enough not to get trarpped by one company's ecosystem over another? For what it's worth, Opus 4.7 has not been a step up and it's come with an enormously higher usage of the subscription Anthropic offers making the entire offering double worse.

I use Open Code as my harness. It's open source, bring your own API Key or OAuth token or self-hosted model. I've jumped from Opus 4.6 to Opus 4.7 to GPT 5.5 in the last 7 days. No big deal, intelligence is just a commodity in 2026. The actual harness is great, very hackable, very extendable.

Does Anthropic not actively ban people using oauth tokens in non-claude-code harnesses?

Re: GPT-5.5

#964
post #305

Earlier quoted context omitted.

I set it up as a joke, to make fun of all of the other benchmarks. To my surprise it ended up being a surprisingly good measure of the quality of the model for other tasks (up to a certain point at least), though I've never seen a convincing argument as to why. I gave a talk about it last year: https://simonwillison.net/2025/Jun/6/six-months-in-llms/ It should not be treated as a serious benchmark.

how can you say "it ended up being a surprisingly good measure of the quality of the model for other tasks" and also "It should not be treated as a serious benchmark" in the same comment? if it is indeed a good measure of the quality of the model (hint: it's not) then, logically, it should be taken seriously. this is, sadly, a great example of the kind of doublethink the "AI" hypesters (yes - whether you like it or n…

I genuinely don't see how those two statements conflict with each other.

Despite not being a serious benchmark (how could it be serious? It's a pelican riding a bicycle!) it still turned out to have some value. You can see that just by scrolling through the archives and watching it improve as the models improved.

If your definition of doublethink is "holding two conflicting ideas in your head at once" then I would say doublethink is a necessary skill for navigating the weird AI era we find ourselves inhabiting.

Re: GPT-5.5

#965

Earlier quoted context omitted.

Did you guys do anything about GPT‘s motivation? I tried to use GPT-5.4 API (at xhigh) for my OpenClaw after the Anthropic Oauthgate, but I just couldn‘t drag it to do its job. I had the most hilarious dialogues along the lines of „You stopped, X would have been next.“ - „Yeah, I‘m sorry, I failed. I should have done X next.“ - „Well, how about you just do it?“ - „Yep, I really should have done it now.“ - “Do X, righ…

This brings up an interesting philosophical point: say we get to AGI... who's to say it won't just be a super smart underachiever-type? "Hey AGI, how's that cure for cancer coming?" "Oh it's done just gotta...formalize it you know. Big rollout and all that..." I would find it divinely funny if we "got there" with AGI and it was just a complete slacker. Hard to justify leaving it on, but too important to turn it off.

Here's a tautology: slacking, consciously refusing to engage agency, requires consciousness and agency. A model can't slack without them.

Re: GPT-5.5

#967
post #529

Earlier quoted context omitted.

We are closer to God than AGI. When AGI arrives, it'll be delivered by Santa Claus.

What do you mean?

It's a multi-layered refute that we are anywhere near AGI while also taking shots at the idea that "God" is real.

And it's taking shots at how far off from Jesus's teachings a lot of "Christianity", particularly those in the media and in power, are..

There is a lot going on there.

Re: GPT-5.5

#968

Earlier quoted context omitted.

Is there any task that actually doesn't require human intervention in-between, even if its just to setup stuff? Like I will get Opus to make me an app but it will stop in between because I need to setup the db and plug in the API keys and Opus really can't do that on its own yet

> Is there any task that actually doesn't require human intervention in-between, even if its just to setup stuff? The goal is none. The current situation: everything that matters requires human intervention. I think the end situation will be that LLMs will be able to perform decently well in a highly controlled and predictable environment.

> in a highly controlled and predictable environment

Why this constraint? A common sentiment I see online (sorry, to group you in) is "[tool] will be capable, actually, but only in a context that trivializes its usefulness."

I think modern post-training like RLVR + inference-time output token scaling can _probably_ scale so the agents can solve any computable task, even when placed in noisy or misconfigured environments. But it won't be economical for a long while. But it already seems largely capable of that today.

Re: GPT-5.5

#969

Earlier quoted context omitted.

This brings up an interesting philosophical point: say we get to AGI... who's to say it won't just be a super smart underachiever-type? "Hey AGI, how's that cure for cancer coming?" "Oh it's done just gotta...formalize it you know. Big rollout and all that..." I would find it divinely funny if we "got there" with AGI and it was just a complete slacker. Hard to justify leaving it on, but too important to turn it off.

Douglas Adams would be proud!

You think you've got problems? What are you supposed to do if you are a manically depressed robot? No, don't try to answer that. I'm fifty thousand times more intelligent than you and even I don't know the answer. It gives me a headache just trying to think down to your level.

Re: GPT-5.5

#970

Earlier quoted context omitted.

LLM models can not do spacial reasoning. I haven't tried with GPT, however, Claude can not solve a Rubik Cube no matter how much I try with prompt engineering. I got Opus 4.6 to get ~70% of the puzzle solved but it got stuck. At $20 a run it prohibitively expensive. The point is if we can prompt an LLM to reason about 3 dimensions, we likely will be able to apply that to math problems which it isn't able to solve cur…

Interesting (would like to hear more), but solving a Rubiks cube would appear to be a poor way to measure spatial understanding or reasoning. Ordinary human spatial intuition lets you think about how to move a tile to a certain location, but not really how to make consistent progress towards a solution; what's needed is knowledge of solution techniques. I'd say what you're measuring is 'perception' rather than reason…

> how to make consistent progress towards a solution

A 7 year old child can learn six sequences of a few moves and over a weekend solve the Rubik Cube. It is a solved algorithm something LLM should be very very good at. What it can't do is reason about spacial relationships.

Post reply on HN