Investors of both should read this: https://open.substack.com/pub/sublius/p/srt-introspect-why-c...
"Investors who have poured hundreds of billions into closed-source labs are betting on an unprovable safety moat".
Nobody is investing in closed-source labs for safety reasons, being able to explore more in details what and how the model is thinking is nice but by no means a game changer. What matters to investors and most of the users is that the model gives the right answer at the end.
In this game, who wins - in the long term - is who has the best model: so far OpenAI is ahead, so in the long term this is what matters. However, for the same reason, if in the future open weight models will be very near the quality of frontier labs, Anthropic and OpenAI will be out of business very soon. The game they play only make sense if their SOTA models do things that other models can't do at a comparable level.
That's a weird way to look at it. Any car gets you to your destination, but some people prefer driving a sports car or an SUV. They get something out of it that isn't just a marketing delusion, but subjective joy from the interaction with one product over another.
Luxury cars are indeed a good comparison. The subjective joy is a result of the delusion. That is why so much money is spent on such marketing to begin with. The analogous comparison would be if a blindfolded passenger turned out to prefer the Sienna to the 911.
At this point I think it’s more important to have a solid workflow and understanding of how [insert your favorite model here] works and its capabilities, than chasing the next shinny release jumping back and forth between companies. I just finished my first large project with Codex and it is hard for me to believe Claude can be much better. It may be a bit better or worse, but again, they are all so good now that the user is the one driving the difference.
That's a weird way to look at it. Any car gets you to your destination, but some people prefer driving a sports car or an SUV. They get something out of it that isn't just a marketing delusion, but subjective joy from the interaction with one product over another.
Luxury cars are indeed a good comparison. The subjective joy is a result of the delusion. That is why so much money is spent on such marketing to begin with. The analogous comparison would be if a blindfolded passenger turned out to prefer the Sienna to the 911.
> The subjective joy is a result of the delusion.
Repeat after me:
_Other people can experience things you do not experience and it is still valid, and not a delusion_. They are not sheeple who fell for marketing.
GPT-5.5 is the better programmer but Opus 4.8 remains the better system architect and product designer. Codex is very "miss the forest for the trees", but is much better at successfully making large changes in large codebases. Claude Code makes more mistakes, but has more taste and a better grasp on idiomatic and elegant software development. If you can afford to, I recommend juggling both.
Great analysis and follows my experience as well. Codex is better when you know how you want the design and the architecture and you drive the agent a lot more aggressively. Claude Code feels like more autopilot so executives and users who didn’t code before AI like it a lot more. But I feel like an expert who can drive GPT aggressively will out perform Opus. It’s why some smart people I know are opting for GPT and h…
This is exactly right. Claude has baked in autonomy and preferences that let it handle underspecified prompts elegantly, which makes it seem smarter to people who like to prompt that way, but it also ignores instructions and fights you on things, which makes it a bad model for people who know what they want to do and specify it.
I never want to hear from developers again that they are not susceptible to marketing. I see meet ups specifically about Claude often. Modern tupperware party. A colleague was convinced Claude is better so we played a game. We used the claude code and codex harness and I implemented some prs they needed with gpt5.5 and opus4.7 and asked them to identify which came from which only from the code. Couldn’t tell. Edit: i…
I don't think it's marketing, for quite a long time Claude was clearly better and not everyone has adapted to the new reality where they have similar capabilities.
I was really frustrated by GPT-5.4, but last night I really pulled out the stops and within a few hours I got path tracing and DLSS implemented on top of Godot, which doesn’t even support DLSS. Just to see if it could do it? And you know what, it did, which was absolutely mind blowing. It wrote like 5,000 lines of C++, I set up a mostly local asset production pipeline using GPT image gen, voiceovers using ElevenLabs API, and even background music using Suno via the chrome use extensions in Codex. I just wanted to see how far I could push this little dumb game my kids asked me to make, and my kids are like “wow our game looks so good!” These models are absolutely mind blowing. I didn’t want to go to sleep I was having so much fun.
It’s because the programming works. OpenAI. Spent its resources on AGI whilst Claude worked on making programming work. Google Gemini is out of the race entirely its programming AI is a joke.
It is unclear which strategy will work in the end. 3.5 flash uses fewer tokens and is cheaper.