Live data from Hacker News

GPT-5.6

openai.com

921–930 of 1001 posts

Re: GPT-5.6

#921
I have Fable send the specs/plans it comes up with to GPT for review and in 2/5 cases yesterday it found additional 1-2 bugs while in the process of reviewing.

GPT-5.6 didn't try to fix the bugs (as instructed) but it did surface them, which is something that didn't happen with GPT-5.5. When spec/plans approved Fable sends back to GPT-5.6 for ralph implementation and it seems to be an even faster, more reliable workhorse than it already was in GPT-5.5. Overall, impressed. Will continue to be a core piece of my workflow.

Re: GPT-5.6

#922
post #72
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

OMP feels like the Linux of AI

Re: GPT-5.6

#923

Earlier quoted context omitted.

> does it really matter anymore? They're different models with different philosophies behind them. This is anecdotal with a user group of 1, but in my experience: Claude has a stronger personality and is more creative. If you give it vague instructions, it's better at filling in the blanks with reasonable ideas. GPT-5.5 is better at following instructions. If you know exactly what you want, it will do it without goin…

I’ve found that Claude is very literal. When I talk to 5.5 it gets what i want it to do, when I talk to Opus 4.8 it does what I say literally and doesn’t get the intent behind it.

This. Claude is very good at instruction and treat my mistake in prompt as law. Gpt is just smarter at understanding my intent.

Re: GPT-5.6

#924
post #862

Earlier quoted context omitted.

We have it slightly ahead of Fable in our multi-agent coding evaluations. Fable's main advantage is that its average solution size is smaller. However, GPT 5.6 Sol is a substantial improvement from GPT 5.4/5.5 which would write verbose, defensive code. 31KB for GPT 5.4/5.5 down to 26KB for GPT 5.6 Sol, with better performance for Sol. Fable scores slightly lower, but with an average solution size of 12.2 KB. Data at…

This looks like a good benchmark. Time and time again I keep giving OpenAI models the chance to win me back, but Opus (and Fable especially) just writes more elegant code and is a significantly more productive rubber duck for interactive discussions. I feel vindicated seeing your description of verbose and defensive code, and I’m a bit disappointed that 5.6 Sol’s solution is still >5x longer than the human solution a…

Really almost all benchmarks I look at have a cost per task column, which is basically the code size metric if you take an extra step

Re: GPT-5.6

#925

GPT-5.6 is the first model where I’ve actually frustrated to use it. I’m explicitly telling it to do something extremely specific and it’s just not listening to me. Eg, I gave it an image to update. The image is sized 400x200 pixels. It then generates a new image at 300x300. I explicitly state to be 400x200 in size and it won’t listen.

Is it possible GPT-5.6 is not a very aligned model?

Re: GPT-5.6

#926
post #893

I just watched the video on their launch page and I am really not sure how I feel about it. On one side, it's cool that these people get to start businesses and stuff using ChatGPT (assuming these are true stories), but how much of the business is really them? And how much does this business rely on a chat bot always being present as a kind of know it all employee? Maybe I'm just being naive or old fashioned (haven't…

You could say the exact same thing about human businesses.

How many ancient Roman businesses are around today?

Re: GPT-5.6

#927
Still fails my internal test of counting legs on animals who have had extra legs photoshopped in. However if prompted to determine what is wrong with the image, it does get it right.

This kind of "out of bounds" image analysis seems to be a very difficult problem to solve, but totally necessary for transformers to really bring about massive change.

Re: GPT-5.6

#928
post #893

I just watched the video on their launch page and I am really not sure how I feel about it. On one side, it's cool that these people get to start businesses and stuff using ChatGPT (assuming these are true stories), but how much of the business is really them? And how much does this business rely on a chat bot always being present as a kind of know it all employee? Maybe I'm just being naive or old fashioned (haven't…

I would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an a…

My code was always plumbing. Glue stuff together. Connect data model to api.

I'm still using my brain, just doing a lot more plumbing now and reading a lot more code than writing. Depressing in some ways and exciting in others. At the end of the day it's not going way. It's disruptive and better to embrace it.

Re: GPT-5.6

#929

Earlier quoted context omitted.

I would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an a…

I think this is where people start to consider you a "Luddite". Is it bad that you used a computer and Google and lifted information from other people to accomplish a task? My father is a machinist and has built has knowledge off the skills and documentation of thousands before him, is that empty? You still have to do the actual work, and where do you draw the line on "shipping more". If a farmer now has an automated…

I am not a luddite by any means, and I think that all these comparisons fall short. Automating physical work has a different effect on our brain than automating intellectual work. Or maybe I'm fooling myself, and it's just our turn as worker to get the industrial revolution treatment.

Re: GPT-5.6

#930
post #38

Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!

The other day I asked Fable about fasting for 16 hours, and it flagged my question. Pathetic situation, this one, where we are supposedly building a superintelligence while at the same time thinking that fasting is a biological weapon.

[deleted]
Post reply on HN