Live data from Hacker News

GPT-5.6

openai.com

931–940 of 1001 posts

Re: GPT-5.6

#931
post #893

I just watched the video on their launch page and I am really not sure how I feel about it. On one side, it's cool that these people get to start businesses and stuff using ChatGPT (assuming these are true stories), but how much of the business is really them? And how much does this business rely on a chat bot always being present as a kind of know it all employee? Maybe I'm just being naive or old fashioned (haven't…

I would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an a…

Those of us in our 50s have been waiting for this day since Eliza.

Re: GPT-5.6

#932
post #891
post #686

Just my two cents. I'm on the Plus plan, I ask gpt-5.6 sol / high to analyze a vibe-coded codebase (~50k LoC) and write a plan to make it production ready. It wasn't a great prompt, I just wanted to test it quickly. It ran for ~15min and consumed 95% of my 5h quota (I thought it was gonna crash). The output is excellent but just a heads up that it consumes a lot of quota!

Anyone looking to use frontier SOTA on the $20 plans is going to have a bad time

If you don't have an agent heavy workflow, you'd be surprised how far the 20$ subscription stretches. I used ~500$ of usage in the last month on a team 20$ plan.

Re: GPT-5.6

#933
It's working great for me.

I generally prefer OpenAI. I usually start projects with Codex; I love the plan we created. Then, after about 2 hours of working, back and forth, etc.I realize it's drifting HARD, or getting stuck on relatively simple things.

Once I get frustrated enough, I sometimes start over with Anthropic. Anthropic has been better (for me at least) to work with from start to finish.

I have the $200/month plan for each of them. (And I used to have the $250/month plan for Gemini... lol?)

Both Fable and Sol are good and definite improvements. I don't have a favorite yet. It's too early to tell.

The biggest difference I'm noticing is usage.

Anthropic is like a used car salesman. Not allowed to do a basic check on your own car (website), they'll scrape every penny possible out of you. Low limits, high api prices, trying to keep everything in their system.

OpenAI is like a cool but aloof dad. Go ahead and borrow his car, shoot, he'll even pay for your gas every once in a while. He'll answer your 5th question drunk, forgetting what you asked. But, at least he's a nice drinker that buys the group shots.

I have to watch my Fable usage, and I'm sure as hell not going to pay API prices. I don't have to watch/worry about my Sol usage even in Ultra.

But, 1 thing I'm noticing using Sol Ultra (on fast mode) is that it's slowing WAY down after a bit. I work the opposite of peak times so that's not it.

Re: GPT-5.6

#934
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

[deleted]

Re: GPT-5.6

#935
post #933

It's working great for me. I generally prefer OpenAI. I usually start projects with Codex; I love the plan we created. Then, after about 2 hours of working, back and forth, etc.I realize it's drifting HARD, or getting stuck on relatively simple things. Once I get frustrated enough, I sometimes start over with Anthropic. Anthropic has been better (for me at least) to work with from start to finish. I have the $200/mon…

May I know if it’s for hobby projects or you run your own business/company where you use subscriptions.

Re: GPT-5.6

#936
post #842

Earlier quoted context omitted.

Given that both Gemini 3.5 Flash (high) and Gemini 3 Flash Preview (medium) beat GPT-5.6 Sol (high) for correctness and score in your benchmarks I don’t trust them at all. The rest of the ranking also doesn’t make sense, like GPT-5.3-Codex (medium) performs better than Claude Opus 4.8 (medium) yeah sure

One example where the order seems correct, is this SVG generation test: https://aibenchy.com/showcase/?q=Gemini+3.5%2Cgpt+5.6%2C+5.3... You can see that most Gemini 3.5 generations are more correct than 5.6 Sol (the net is in the middle of the table, hamster seems reasonable and not deformed, etc.)

It doesn't seem correct at all though? the supposedly best one isn't even the best gemini flash output (the medium one looks better than the high one)

Re: GPT-5.6

#937
post #97

Earlier quoted context omitted.

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…

Honestly it’s the usage limits that are so generous that makes codex worth it even if it may not be exactly as powerful as Claude. The peace of mind that you can try a lot of things and make huge refactors and run extensive redundant tests without running out of tokens just makes the whole thing a much better experience. I tried coding with Deepseek and it was pretty terrible so the only reason codex works is because…

> I tried coding with Deepseek

But soooooo cheap. Especially for those of us where a monthly sub doesn’t make sense.

Re: GPT-5.6

#938

Earlier quoted context omitted.

I would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an a…

I also feel this, and it also troubles me. Unfortunately, we may be in the minority on this site. What the other replies seem to overlook is that it fundamentally changes the nature of the work - it’s not the next step in the evolution after an IDE, it is closer to an automated tractor, and yes that does make the farmer’s work trivial. Pressing a button and having the field get plowed is a very different experience t…

I am on the side that says the farmer didn't plow the field. I guess you could say he took responsibility for it, which isn't the same thing. I guess a counter argument that some would make is that the farmer didn't plow the field if they didn't pull the plow themselves (or to a silly level: if they didn't plow the field with their bare hands). They just pressed the right pedals while sitting in the tractor. I am not sure I have a good argument either way, but a farmer plowing a field with a tractor is physically involved during the whole process, and it feels more fair for them to say that they plowed the field, for the simple fact they weren't doing anything else during the process other than constantly controlling the plowing machinery.

Re: GPT-5.6

#939
post #933

It's working great for me. I generally prefer OpenAI. I usually start projects with Codex; I love the plan we created. Then, after about 2 hours of working, back and forth, etc.I realize it's drifting HARD, or getting stuck on relatively simple things. Once I get frustrated enough, I sometimes start over with Anthropic. Anthropic has been better (for me at least) to work with from start to finish. I have the $200/mon…

May I know if it’s for hobby projects or you run your own business/company where you use subscriptions.

[deleted]

Re: GPT-5.6

#940
post #933

It's working great for me. I generally prefer OpenAI. I usually start projects with Codex; I love the plan we created. Then, after about 2 hours of working, back and forth, etc.I realize it's drifting HARD, or getting stuck on relatively simple things. Once I get frustrated enough, I sometimes start over with Anthropic. Anthropic has been better (for me at least) to work with from start to finish. I have the $200/mon…

May I know if it’s for hobby projects or you run your own business/company where you use subscriptions.

[deleted]
Post reply on HN