GPT-5.6 didn't try to fix the bugs (as instructed) but it did surface them, which is something that didn't happen with GPT-5.5. When spec/plans approved Fable sends back to GPT-5.6 for ralph implementation and it seems to be an even faster, more reliable workhorse than it already was in GPT-5.5. Overall, impressed. Will continue to be a core piece of my workflow.
GPT-5.6
921–930 of 1001 posts
Re: GPT-5.6
#922Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right
Re: GPT-5.6
#923Earlier quoted context omitted.
> does it really matter anymore? They're different models with different philosophies behind them. This is anecdotal with a user group of 1, but in my experience: Claude has a stronger personality and is more creative. If you give it vague instructions, it's better at filling in the blanks with reasonable ideas. GPT-5.5 is better at following instructions. If you know exactly what you want, it will do it without goin…
I’ve found that Claude is very literal. When I talk to 5.5 it gets what i want it to do, when I talk to Opus 4.8 it does what I say literally and doesn’t get the intent behind it.
Re: GPT-5.6
#924Earlier quoted context omitted.
We have it slightly ahead of Fable in our multi-agent coding evaluations. Fable's main advantage is that its average solution size is smaller. However, GPT 5.6 Sol is a substantial improvement from GPT 5.4/5.5 which would write verbose, defensive code. 31KB for GPT 5.4/5.5 down to 26KB for GPT 5.6 Sol, with better performance for Sol. Fable scores slightly lower, but with an average solution size of 12.2 KB. Data at…
This looks like a good benchmark. Time and time again I keep giving OpenAI models the chance to win me back, but Opus (and Fable especially) just writes more elegant code and is a significantly more productive rubber duck for interactive discussions. I feel vindicated seeing your description of verbose and defensive code, and I’m a bit disappointed that 5.6 Sol’s solution is still >5x longer than the human solution a…
Re: GPT-5.6
#925GPT-5.6 is the first model where I’ve actually frustrated to use it. I’m explicitly telling it to do something extremely specific and it’s just not listening to me. Eg, I gave it an image to update. The image is sized 400x200 pixels. It then generates a new image at 300x300. I explicitly state to be 400x200 in size and it won’t listen.
Re: GPT-5.6
#926I just watched the video on their launch page and I am really not sure how I feel about it. On one side, it's cool that these people get to start businesses and stuff using ChatGPT (assuming these are true stories), but how much of the business is really them? And how much does this business rely on a chat bot always being present as a kind of know it all employee? Maybe I'm just being naive or old fashioned (haven't…
How many ancient Roman businesses are around today?
Re: GPT-5.6
#927This kind of "out of bounds" image analysis seems to be a very difficult problem to solve, but totally necessary for transformers to really bring about massive change.
Re: GPT-5.6
#928I just watched the video on their launch page and I am really not sure how I feel about it. On one side, it's cool that these people get to start businesses and stuff using ChatGPT (assuming these are true stories), but how much of the business is really them? And how much does this business rely on a chat bot always being present as a kind of know it all employee? Maybe I'm just being naive or old fashioned (haven't…
I would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an a…
I'm still using my brain, just doing a lot more plumbing now and reading a lot more code than writing. Depressing in some ways and exciting in others. At the end of the day it's not going way. It's disruptive and better to embrace it.
Re: GPT-5.6
#929Earlier quoted context omitted.
I would add to this that, to me and to many friends in their 30-40s, using AI models to achieve something we used our brains to achieve feels... empty? wrong? soulless? Sure, a lot of menial work can be relegated to the models, and it's ok, most of the time, but you finish a day of work with the inability to shake a precise new feeling: that you haven't really achieved something, even if you shipped more than on an a…
I think this is where people start to consider you a "Luddite". Is it bad that you used a computer and Google and lifted information from other people to accomplish a task? My father is a machinist and has built has knowledge off the skills and documentation of thousands before him, is that empty? You still have to do the actual work, and where do you draw the line on "shipping more". If a farmer now has an automated…
Re: GPT-5.6
#930Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!
The other day I asked Fable about fasting for 16 hours, and it flagged my question. Pathetic situation, this one, where we are supposedly building a superintelligence while at the same time thinking that fasting is a biological weapon.